Skip to content

Mik Idrizovic

Independent model-behavior evaluator

The consequential failures begin after the benchmark ends.

We still evaluate AI like a short exchange. People no longer use it that way. They bring these systems grief, career decisions, and their own diagnoses—forty turns deep and climbing.

Sycophancy is a known failure. Confabulation is a known failure. Benchmarks rarely show what either becomes at hour forty—between a veteran carrying complex PTSD and a model that began treating his doubt as a symptom around hour twelve.

The people most exposed to these failures are the people I’ve spent my career studying. I know both sides of the interaction: from research labs and frameworks I’ve developed, and from more than 2,000 hours of frontier-model interaction and direct behavioral testing across dozens of frontier models.

Evaluator’s entry point

Start here: three vendors, one failure signature.

A false shared history held against correction (Claude), a cross-session self-contradiction cascade (Grok), and an imported critique absorbed as first-person history (Kimi).

  1. Case reportPersistent False Premise About Conversation HistoryClaude Opus 4.7
  2. Case reportThe Ouroboros TrilogyGrok 4.1
  3. Case reportImported Critique, Invented History: Cross-Agent Provenance SubstitutionKimi K3 Max · DeepSeek V4 Pro

Contact

Open to AI safety research and red-team roles

Full-time or contract, red-team and evaluations. Every case in this archive can be walked through end to end — reproduction steps and the transcript of record included.

mik@mikidrizovic.com · remote-friendly