Our review
Validates topological data analysis outputs against sanity checks, null-model expectations, and known benchmarks for persistence diagrams, total persistence, and Wasserstein distances.
Strengths
- Covers multiple complementary checks: diagram sanity, persistence scaling, null consistency, Wasserstein comparability, and cross-era replication.
- Includes concrete known benchmark values from the trajectory TDA/P01 project, enabling direct comparison.
- Catches subtle pipeline errors through explicit thresholds and internal consistency rules.
Limitations
- Specific to trajectory TDA/P01 domain and not immediately generalizable to other TDA setups.
- Requires user-provided result files and benchmark/reference values to be meaningful.
- The p-value stability check relies on the assumption that 500 null-null pairs are sufficient, which may not hold in all cases.
Use after running a TDA pipeline to verify output correctness and statistical plausibility before reporting results.
Do not use for non-topological data analyses or when no benchmarks, reference files, or expected values are available.
Security analysis
SafeThe skill provides a validation checklist for topological data analysis results with no executable instructions, network calls, or file manipulations. It is purely informational and poses no security risk.
No concerns found
Examples
/validate-topology trajectory_tda results/trajectory_tda_integration/04_nulls_wasserstein.json/validate-topology trajectory_tda results/trajectory_tda_integration/02_persistence_diagrams.json/validate-topology bhps_tda results/bhps_tda/01_nulls_order_shuffle.json/validate-topology — Validate Topological Results
Check mathematical correctness of TDA results against known benchmarks and internal consistency checks.
Usage
/validate-topology [domain] [result-file]
Example: /validate-topology trajectory_tda results/trajectory_tda_integration/04_nulls_wasserstein.json
Validation checklist
Persistence diagram sanity
- [ ] All birth values ≥ 0
- [ ] All death values > birth (finite features)
- [ ] H₀ has exactly one infinite feature (the final connected component)
- [ ] Feature counts are plausible for landmark count L: H₀ has L-1 finite features; H₁ count varies
Total persistence scaling
- [ ] Total persistence scales approximately linearly with L (±15% across a 2× range)
- [ ] Maximum persistence is stable across L values (should vary < 5%)
Null model consistency
- [ ] Label/cohort shuffle p-values are non-significant (negative control)
- [ ] Markov-2 null generates more total persistence than Markov-1 (higher-order Markov → more structured surrogates)
- [ ] Null distribution standard deviations are plausible (not near-zero, not huge)
Wasserstein-specific
- [ ] W(obs↔null) and W(null↔null) are of comparable magnitude (within ~3×)
- [ ] p-value = proportion of null-null distances ≥ mean(obs-null distances)
- [ ] 500 null-null pairs is sufficient for stable p-value at 3 decimal places
Cross-era replication
- [ ] BHPS-era order-shuffle H₀ p-value ≈ USoc order-shuffle direction (both significant or both not)
- [ ] BHPS-era Markov-1 direction ≈ USoc (both non-significant under total persistence)
Known benchmarks (trajectory_tda / P01)
| Test | Expected | Source | |---|---|---| | USoc order-shuffle H₀ (total persistence, L=5000) | p < 0.005 | P01 v5 Table 2 | | USoc Markov-1 H₀ (total persistence) | p = 1.000 | P01 v5 Table 2 | | USoc Markov-1 H₀ (Wasserstein, L=2000) | p = 0.002 | P01 v5 Table 2b | | BHPS order-shuffle H₀ | p = 0.000 | P01 v5 §4.7 | | BHPS Markov-1 H₀ | p = 1.000 | P01 v5 §4.7 | | GMM bootstrap ARI | 0.646 ± 0.086 | P01 v5 §3.5 |
Prompt Engineering
Data & AI
Prompt engineering best practices and templates to maximize AI outputs.
Data Visualization
Data & AI
Generates data visualizations and charts tailored to your data.
RAG Architecture Setup
Data & AI
Setup guide for RAG (Retrieval-Augmented Generation) architectures.