From Averages to the Whole Pyramid: Mapping Every Tier of Evidence with the Avinash Principle
The problem with stopping at case series
The last piece in this series, From Averages to Trajectories: Mapping Case Series with the Avinash Principle, showed something simple but useful: run the same three-step pipeline — extract nodes, merge into a network, render an interactive plot — across a whole case series, and the textbook path and outlier forks become visible as one picture instead of a table of aggregate percentages.
But a case series is only one rung of the evidence pyramid. It's the rung with the least structure to lean on — no comparison group, no randomization, no pooling. The moment you try to point the same mapper at a case-control study, a cohort study, an RCT, or a meta-analysis, it breaks, because the unit of analysis changes and the direction of inference changes with it. A case series has nothing to compare against. A case-control study looks backward from a shared outcome. A cohort study looks forward from a shared exposure. A cross-sectional study has no time axis at all. An RCT has a CONSORT flow the paper is required to report. A meta-analysis isn't built from patients at all — it's built from studies.
Treat all of these the same way, and the tool either fabricates trajectories the data can't support, or draws visual relationships — a left-to-right arrow, a single "textbook path" — that quietly imply more than the study design was ever built to prove. That's the seam the Evidence-Pyramid Trajectory Mapper picks up: one router, six modules underneath it (one per design tier), plus a seventh combining mode for when several papers on the same question are uploaded together and you want to watch the signal survive — or not — as it climbs from anecdote to meta-analysis.
Step 0: classify before you visualize
Nothing runs before the paper is classified. The prompt reads the title, abstract, methods, and whatever reporting checklist the paper cites — CONSORT, PRISMA, STROBE, CARE — and sorts it into exactly one of six primary designs: case report/series, case-control, cohort, cross-sectional, RCT, or systematic review/meta-analysis. Then it states that classification and the specific evidence for it, out loud, before doing anything else.
This matters because papers are frequently mixed. An RCT with an embedded case series of adverse events. A meta-analysis with a worked example. The rule here is the same discipline as the earlier pieces in this series: name the dominant design, note that a secondary module could apply to the embedded sub-analysis, but build only one primary visualization unless asked for both. Two incompatible visual grammars blended into a single file is worse than one honest, partial file. If the classification is genuinely ambiguous, the prompt states the two most likely designs, explains what in the paper points to each, and proceeds with the better-supported one rather than stalling.
Once classified, everything downstream shares the same skeleton — EXTRACT → MERGE → VISUALIZE — but the unit of analysis and the node/edge semantics change with the design.
Six modules, six units of analysis
Module A, Case Report/Series (the Trajectory Mapper) is the one this series already covered — one patient, one trajectory, 3–6 nodes running from presentation through decision/intervention through reaction/complication to disposition. Hub nodes, a textbook path, and outlier forks, rendered as a sized-and-weighted node-link diagram.
Module B, Case-Control (the Divergence Mapper), runs the same idea backward. Each subject's chain is traced from a shared outcome anchor back through exposure history to confounders and how exposure was actually ascertained. Cases and controls become two mirrored sub-networks sharing the same node schema so they can be overlaid, and the object of interest isn't a single textbook path but the divergence path — the node where case-density and control-density pull apart most sharply, annotated with the paper's own odds ratio and CI if reported. The hard constraint here is specific and easy to violate by accident: an odds ratio is never allowed to be silently reframed as a "risk" or "probability" — case-control designs simply don't estimate absolute risk, and the visualization has to say so.
Module C, Cohort (the Forward Trajectory / Incidence Mapper), runs forward instead — baseline exposure status, follow-up checkpoints with elapsed time, an outcome node, and a censoring node that is explicitly not treated as "no event." Cohort studies naturally produce two textbook paths, one exposed and one unexposed, rather than one — forcing them into a single line would erase the comparison the whole design exists to make. The divergence point is the follow-up checkpoint where the two incidence curves separate most, annotated with whatever relative risk or hazard ratio the paper reports — never one computed by the tool itself.
Module D, Cross-Sectional (the Snapshot Association Mapper), is the one design with no time axis at all, and the module is built to make that limitation impossible to miss. Instead of sequential nodes it uses association layers — demographic stratum, exposure/status, outcome/condition, and reported association, all measured at the same instant. The network itself becomes a co-occurrence graph rather than a path: nodes are strata, edges connect categories that co-occur, and there is no textbook path in the sequential sense at all. The non-negotiable constraint: never draw an arrow or any left-to-right implication between exposure and outcome nodes, because a cross-sectional design cannot support directionality and the picture must not visually suggest otherwise.
Module E, RCT (the Arm-Comparison / CONSORT Trajectory Mapper), works at two layers simultaneously. Layer one is always extractable from an RCT by definition — the CONSORT flow itself, assessed-for-eligibility through randomized, allocated, and analyzed, per arm, using the paper's actual reported numbers at every stage (never estimated by subtraction across mismatched tables). Layer two is patient-level detail, but only where the paper actually reports it beyond the flow diagram; where it doesn't, the module falls back honestly to arm-level summary nodes rather than simulating individual patients to make the diagram look richer than the evidence is.
Module F, Systematic Review/Meta-Analysis (the Study-Level Trajectory / Forest Mapper), is where the unit of analysis itself moves up a level: one included study is one trajectory, not one patient. Each study in the review's evidence table gets a short chain — identification, population/setting, exactly how the intervention or exposure was defined (definitional drift across studies is often the real heterogeneity story), the study's own effect estimate and analytic weight, and its risk-of-bias rating if the review assigned one. The visualization links a forest plot directly to a node network of the shared population/definition/design categories each study passes through — click a study in the forest plot, its path through the network highlights, and vice versa. The constraint that matters most here: never pool an effect estimate yourself. If the review is narrative/qualitative rather than a formal meta-analysis, no forest plot and no pooled diamond get drawn at all — the module drops back to a study-level co-occurrence network instead, the same honest downgrade Module D uses for cross-sectional data.
Module G: watching the signal climb
The payoff of building six separate modules is that they can be stacked. Module G, the Evidence-Pyramid Climb, activates only when more than one paper on the same underlying question — spanning at least two design tiers — has been uploaded, and it runs after each paper has already gone through its own Module A–F individually. It never replaces the individual analysis; it sits on top of it.
Before anything gets combined, G0 forces an honest comparability check. Do the papers actually share a comparable population/exposure-or-intervention/comparator/outcome frame, on roughly the same timeframe? If one paper measures symptom relief and another measures mortality, those are different constructs, and the module says so rather than forcing them into one view. Papers get ranked by their own already-established pyramid tier, and anything that can't be reasonably compared — wrong population, wrong outcome, wrong timeframe — gets explicitly excluded with a stated reason, rather than silently dropped or silently forced in.
What survives that filter gets distilled (G1) into a comparable signal chain per paper — tier, population/frame, effect-and-direction (kept in its own native units: a case series' qualitative impression is never converted into a fabricated numeric effect size, and an RCT's confidence interval is never flattened into an anecdote), and a certainty/bias rating if the paper itself reported one.
G2 is where the actual structure appears: a vertical, tiered pyramid — case reports and series at the base, meta-analyses at the apex — with each paper placed as a node in its tier's band, sized by its own native weight. Vertical edges connect papers in adjacent tiers addressing the same question, and each edge gets classified as reinforcing (direction and magnitude agree as the tier climbs), attenuating (same direction, but the effect shrinks substantially at the higher tier — often the first sign a lower-tier signal was partly confounded), contradicting (direction reverses), or unresolved/novel (a signal exists at only one tier, with no paper yet addressing it above or below — and the module flags whether that's because no one has studied it, or simply because it wasn't given that paper to read). The climb path — the single clearest evidentiary chain running bottom to top — gets highlighted distinctly, and any point where a lower-tier signal is directly contradicted by a higher-tier design is marked as the single most important node in the whole diagram, never buried beside routine ones.
The constraint on G is as strict as anywhere else in the system: never compute a pooled or averaged effect across papers unless one of the included papers is itself a meta-analysis already reporting that pooled estimate. This module visualizes concordance across tiers — it does not perform new statistical synthesis of its own.
What stays constant across all seven modules
A few requirements don't change no matter which module fires, and they're worth naming because they're the difference between a tool that's honest about evidence and one that just looks tidy.
Every output — regardless of module — follows the same order: the design classification and its supporting evidence, an extraction summary stating how many units had extractable trajectories and which nodes turned out to be hubs versus outlier-only, a plain-text hub/textbook-path (or pooled-estimate/divergence-path) summary, a plain-text outlier/divergence summary naming the most informative divergences and why each looks meaningful rather than noise, the interactive HTML file itself, and — mandatory, every time, final section — a reference list.
That reference requirement runs deeper than a bibliography tacked on at the end. Every named study, patient, outlier, or pooled-estimate claim in the chat-facing summary has to carry an inline citation back to its source — a paper section, a table or figure number, or, for Modules F and G where the units of analysis are themselves cited studies, the original study's own reference-list citation. Inside the HTML file itself, every node that corresponds to a specific cited source has to show that citation in its hover tooltip and in any visible sources panel — an anonymous "Study 7" label is not acceptable when the paper gives an author, year, or reference number. And if a node genuinely synthesizes several studies with no single attributable source, it gets labeled explicitly as a derived/aggregate node rather than assigned one invented citation.
Visually, every module produces output in the same Vibe Rounds house style as the rest of the suite — the teal/cyan accent for hub and textbook-path elements, the orange/rust alert color reserved for outliers, contradictions, and high-risk-of-bias flags, Inter for body text and a monospace face for labels and data values, and every supporting panel (the network detail, the outlier breakdown, the forest-plot panel, the sources list) collapsed by default so the page loads clean and the reader opens only what they want to inspect. A disclaimer banner sits above all of it, stating plainly that this is an educational tool for reading evidence structure, not a diagnostic instrument.
The same discipline, one level up each time
The through-line connecting every piece in this series is a refusal to let a picture claim more than the underlying data actually supports. The first piece showed that expert cognition prunes a quarter-billion theoretical trajectories down to the handful of hub nodes that actually carry a case's gravity. The case-series piece asked the same question across a cohort: where do trajectories genuinely overlap into a real textbook path, and where does thin reporting quietly stop telling the truth about any one patient? The Evidence-Pyramid Mapper asks it at the highest level available — across an entire body of literature, from a single anecdote up to a formal meta-analysis — where does a signal actually hold as the evidence gets stronger, and where does it fall apart the moment someone looks harder?
Six modules for six designs, one combining layer for when several of them are stacked against the same question, and one rule that never changes regardless of which tier you're standing on: never let a visualization imply a stronger causal, temporal, or evidentiary claim than the study design underneath it was ever built to prove.
This piece is part of a series on the Avinash Principle and the Vibe Rounds tools built from it. Earlier pieces: The Avinash Principle: Critical Hub-Node Navigation, From Trajectory Analysis to Trajectory Prediction and Control, From Averages to Trajectories: Mapping Case Series with the Avinash Principle, and the synthesis, From One Trajectory to Every Trajectory. None of this — including this piece — is meant as ground truth about any underlying disease or as a substitute for clinical judgment. It's scaffolding for a learnable skill: finding which handful of nodes actually carry a trajectory's gravity, and being honest about what the evidence can and can't tell you at every level, from one patient up to an entire literature.
No comments:
Post a Comment