Methods and Frameworks
The CMD Research Group takes a distinctly mathematical approach to questions that might otherwise be treated purely empirically. We believe that without formal frameworks, empirical observations about AI systems remain uninterpretable – interesting patterns without explanatory power.
Ollivier-Ricci curvature, graph topology, and embedding space analysis provide the mathematical substrate for our core theoretical claims.
Propensity score matching and related methods allow us to make causal claims about drift, not merely correlational observations.
BioBERT embeddings, transformer probing, and diachronic corpus analysis provide the empirical grounding for our theoretical claims.
We treat semantic drift as a dynamical system phenomenon. Meaning states evolve according to transformation rules imposed by AI processing. Attractors, basins, and bifurcation points in this system correspond to stable meanings, regions of semantic ambiguity, and points at which small perturbations produce qualitatively different outputs.
This framing has practical consequences. It means drift is not simply random noise – it has structure. And structured phenomena can be predicted, controlled, and exploited.
EchoCodes is our flagship framework for detecting semantic drift in clinical diagnostic coding. It applies BioBERT embeddings to ICD-10 code descriptions across time and processing stages, using geometric metrics to flag codes whose semantic neighborhoods have shifted beyond a calibrated threshold.
The framework is designed for deployment in large health systems and produces surveillance-grade alerts when diagnostic code drift exceeds clinically meaningful levels.
Our work targets collaboration with DoD and Army research organizations including ARO and ARL, with additional applications in healthcare informatics and civilian AI safety.
We welcome inquiries from researchers, practitioners, and funding partners working in adjacent areas. See the Contact section of the main site.