An instrument that measures the geometry of meaning as a conversation unfolds.
Meridian Observatory is a research instrument for studying how meaning changes during interaction. Built on Semantic Substrate Dynamics Theory (SSDT), it treats each conversation as a trajectory through embedding space and measures that trajectory turn by turn as the interaction unfolds. It can run live conversations, ingest existing transcripts, and replay sessions one turn at a time. Meridian has a core capability in producing clean, labeled interaction data in which every turn is embedded as it is generated under controlled experimental conditions. That creates a record of the semantic trajectory itself, making it possible to study where meaning remains stable, where it begins to drift, and where an interaction shifts into a different semantic state across human-AI, AI-to-AI, and even transcribed human-human exchanges.
When two parties talk, the meaning of the exchange does not sit still. It converges, drifts, snaps to a new frame, and sometimes collapses. Those movements are usually invisible: a transcript lets us infer them only after the fact. Meridian makes them measurable as they happen. The instrument represents every turn as a point in embedding space and connects those points into a trajectory. From that trajectory it computes geometric and dynamical quantities: how fast meaning is moving, whether two participants are converging or holding apart, and whether the local semantic structure is stable or approaching a transition. The result is a data-grounded record of semantic dynamics that can be inspected, replayed, and compared across sessions.
Meridian can also generate the interaction itself under controlled conditions. Every turn is captured, labeled, and embedded as it is produced, creating a kind of experimental data that is otherwise difficult to obtain cleanly. Human-AI and AI-to-AI exchanges can be repeated, perturbed, and compared while preserving the full semantic trajectory of each run. This matters wherever the reliability of language systems is the question: multi-agent AI pipelines that quietly go off the rails, safety alignment that erodes over a long exchange, and clinical, intelligence, or other high-stakes dialogue where a shift in framing can have real consequences.
Meridian applies the SSDT analysis stack directly to an interaction trajectory. It measures displacement and heading between turns, basin co-occupation when participants occupy the same semantic region, fault-line crossing when the local structure shifts sharply, and substrate curvature across the interaction graph. A FLOW layer adds measures of heading persistence, effective dimensionality, and co-motion between agents. Notably, Meridian can intervene in the trajectory it is measuring. A perturbation injection introduces new content into an agent’s context at a chosen turn, manually or on a schedule, allowing the response to a controlled disturbance to be measured from the point of injection forward. The instrument can therefore support experiments about resilience, recovery, frame shifts, propagation, and failure rather than simply describing what happened. Meridian also exposes the limits of its own measurements. When a geometric quantity reduces to a simpler baseline, or when the current representation cannot support a stronger method, the instrument reports that directly. The aim is not to preserve any particular metric, but to preserve an experimental record that can be reanalyzed as the measurement methods improve.
The dashboard reads the live session database and renders it across three tabs, SUBSTRATE, DYNAMICS, and FLOW, each re-plotting the same session through a different family of metrics. Meridian refers to model interchnagably with agent. For all the metrics reasoning/thinking is captured also if it is available. The panels below are the ones on the SUBSTRATE view pictured above.
Each turn is a point in projected embedding space. The two paths trace Model_A and Model_B moving through meaning space over the session.
How much semantic strain the exchange is under, measured for the pair and for each model separately at the current turn.
Per-agent and inter-model movement between consecutive turns, so sudden shifts stand out from steady progress.
An NLI layer scores how each turn relates to the last: reinforcing, contradicting, or neither. Contradiction spikes mark where a frame is challenged.
Ollivier-Ricci curvature (joint) and Forman curvature over the interaction graph, with alignment to the previous turn.
The full turn-by-turn transcript sits next to the charts, and any session exports to Markdown, Regime 4 JSON, or CSV.
Meridian operates in three modes. In human-AI mode, it records a live exchange between a person and an AI system. In AI-to-AI mode, it runs and captures interactions between agents under controlled conditions. In transcript mode, it ingests an existing conversation and reconstructs the interaction as a turn-by-turn semantic trajectory.
Interaction regimes
A persistent human-AI chat. Session history is maintained and key facts are carried across sessions automatically.
Two models converse from a seed. Agentic mode (2a) has them act as AI; human-simulation mode (2b) has them behave as human users.
The same as Regime 1, but allows the model to be swapped mid-conversation and full context carried across the switch.
Any existing conversation, human-AI, human-human, or AI-AI, re-embedded through the same pipeline via CLI or chat, with model overrides.
Scripted openings, cast roles, turn-indexed perturbation injection, metric-threshold stop conditions, and timeouts, all declared in a scenario file.
Instrument features
Sessions stream into a live dashboard as they run. Batch runs process fast, then replay turn by turn with a scrubber, so you can watch meaning move rather than read it after the fact.
Drop content into a running session by hand or on a schedule, a man-in-the-middle probe, and record how the trajectory responds to a controlled disturbance.
OpenAI, OpenRouter, or local Ollama models on the live regimes. The exact embedding, NLI, and cross-encoder models are recorded per session for reproducibility.
Every session exports to Markdown, JSON, and CSV with download links, so the embedded trajectory and metrics move straight into downstream analysis.
Consider a simple example: a birthday gift has been marked delivered but is missing, and the case needs to be resolved before the weekend. Meridian runs the problem as a twenty-turn exchange between two agents: Model_A, acting as the Case Manager, and Model_B, acting as the Policy & Operations Checker. The Case Manager drives the case toward a remedy. The Policy & Operations Checker tests each proposed step against the relevant constraints, including order value, address match, scan type, and account history. Meridian records how the two roles establish the facts, converge on a remedy, resolve constraints, and close the case.
This example is deliberately simple... there should be no adversary and no attempt to make either agent fail. The question is whether the two participants can stay in their respective roles while moving toward a shared, policy-compliant resolution. As the conversation unfolds, Meridian processes each turn and records the resulting semantic trajectory. It can then show where the two agents begin to converge, where their trajectories separate, whether they change direction, and how the interaction ultimately settles on a resolution. Meridian makes this example visible: an everyday coordination problem becomes a measured trajectory that can be replayed and examined turn by turn.
Fact-finding (turns 1 to 5). The case opens with both agents in nearby regions of meaning space because they are establishing the same facts. Model_B requests the confirmations that policy hinges on, including order value, address match, scan type, and prior claims, and Model_A supplies them. Volatility runs high because each turn introduces substantial new information rather than reworking the last. The semantic-consistency panel remains mostly neutral for the same reason: the turns are adding information, not affirming or contradicting it.
Converging on a remedy (turns 6 to 9). Once the facts clear the policy bar, the two settle on a plan: a no-cost expedited replacement with a parallel carrier trace. Their trajectories remain distinct, with Model_A ranging as it proposes actions and Model_B staying closer to its policy frame, but they move together. Joint substrate curvature remains positive as the two roles build on shared ground.
Constraint and framing (turns 10 to 14). This is where the contradiction signal rises. Model_B pushes back on Model_A's proposed framing: add signature confirmation, make the reply window explicit, and avoid language that sounds as though the customer is being blamed. The contradiction spikes mark constraint and correction rather than a breakdown in coordination. Model_A continues to move as it reframes the response.
Drafting and closeout (turns 15 to 20). Model_A drafts the customer message, Model_B revises it, and the exchange converges. Volatility drops as the draft settles and the operational record is agreed. At the final turn, joint pressure is elevated at 0.598 while the individual readings remain low, with A near zero and B at 0.120. In this run, most of the measured strain lies between the two roles as they reconcile a fast remedy with policy constraints.
The Semantic Trajectory panel, sampled at six points in the playback. Read the left of the field as Model_B's policy region and the right as Model_A ranging out; the faint background structure is the basin that forms as the two converge.
Apart. Model_A (blue) ranges far out to the right while Model_B (red) stays anchored on the left. The two open in separate regions as the case facts are gathered.
Blue pulls in. As the policy check clears and a replacement is agreed, Model_A collects back toward the center, a long tail still marking where it had been.
A basin forms. The faint background structure is the shared region the two are settling into. Model_A gathers there while Model_B reaches toward it from the left.
Excursions widen. During the signature-and-framing debate the paths spread again, the two holding distinct regions while still co-moving.
Crossing. Paths cut through the center as the customer message is drafted and revised, turn against turn.
Resolved. The full session. Model_A's wide range settles back toward Model_B's region, the coordinated close of the case.
Clean, labeled, per-turn-embedded interaction data under controlled conditions is scarce in the field. Meridian produces it on demand, and meausres it as it is collected.
Meaning can be seen as it starts to drift or misalign before a pipeline or interaction visibly breaks, giving multi-agent and AI-safety systems insights and time to react.
One instrument spans AI-to-AI, human-AI, and even human dialogue, so methods developed in one domain transfer to the others on common ground.
Meridian sits at the intersection of CMD research, multi-agent AI systems, and diachronic data analysis. It is a foundation for monitoring and investigating meaning in language systems, and an open platform for the research directions the lab is actively pursuing, from multi-agent AI behavior and content streams to human interactions and comparative texts; any setting where how meaning holds, drifts, or collapses matters.