Independent R&D · 2026
Measuring whether a persona survives real use
A reproducible evaluation program for identity, voice, conversational continuity, and drift across changing models and longer interactions. The public story is about measurement design—not the underlying product implementation.
- Scope
- Multi-model evaluation, attribution, drift, and conversational-flow checks
- Discipline
- Fixed cohorts, honest baselines, confidence intervals, and reproducible reports
- Result
- Found the operating regime where pooled evidence is useful—and documented where it is not
What stays private: product-specific prompts, live infrastructure, customer workflows, private datasets, and deployment configuration.