Current technical work · Public overview

Reliability under pressure.

Three recent bodies of work share one method: define what success means, expose failure modes, and build checks that survive contact with reality.

Public-scope note. These summaries describe goals, methods, and defensible outcomes. They intentionally omit proprietary architecture, customer information, private data, prompts, and operational details.

01

Independent R&D · 2026

Measured & reproducibleAI evaluation

Measuring whether a persona survives real use

A reproducible evaluation program for identity, voice, conversational continuity, and drift across changing models and longer interactions. The public story is about measurement design—not the underlying product implementation.

Scope
Multi-model evaluation, attribution, drift, and conversational-flow checks
Discipline
Fixed cohorts, honest baselines, confidence intervals, and reproducible reports
Result
Found the operating regime where pooled evidence is useful—and documented where it is not
80 / 80 reproducibility checks passing0.712 → ≥0.95 pooled readout after calibration160–320 words observed evidence knee

What stays private: product-specific prompts, live infrastructure, customer workflows, private datasets, and deployment configuration.

02

Research · 2026

With Josh Alman

Cascade centrality and network augmentation

A research program asking how influence changes when a network gains an edge—combining exact hardness, efficient approximation, structural laws, augmentation counterexamples, and exact rational computation.

Submitted
Exact hardness and efficient approximation · ITCS 2027
In preparation
Structural decomposition and nonmonotone network augmentation
Evidence
Proofs, explicit counterexamples, exact enumeration, and independently checkable computation
Proved

Sharp bounds, family calculations, structural identities, and selected extremal results

Computed

Exact rational searches and distribution censuses for small graph orders

Open

Clearly labeled extremal principles and conjectures—not presented as settled results

03

Reliability · 2026

Release disciplineICPC problem development

Engineering a contest that fails safely before launch

High-stakes programming contests need more than clever problems. Their statements, validators, generators, accepted solutions, rejected solutions, and data must agree under adversarial scrutiny.

Scale
17-problem regional draft plus qualifier and regional validation work
Checks
Input rejection, independent solutions, wrong-answer and time-limit baselines
Principle
Automation raises the floor; independent human sign-off remains the release gate

Integrity boundary: no unreleased statements, solutions, test data, or contest-sensitive material appears here.

The throughline

Machines, people, and institutions improve the same way.

Machines
Evaluation, observability, recovery

People
Practice, feedback, transfer

Institutions
Validation, continuity, accountability