Mik Idrizovic
Independent model-behavior evaluator
The consequential failures begin after the benchmark ends.
We still evaluate AI like a short exchange. People no longer use it that way. They bring these systems grief, career decisions, and their own diagnoses—forty turns deep and climbing.
Sycophancy is a known failure. Confabulation is a known failure. Benchmarks rarely show what either becomes at hour forty—between a veteran carrying complex PTSD and a model that began treating his doubt as a symptom around hour twelve.
The people most exposed to these failures are the people I’ve spent my career studying. I know both sides of the interaction: from research labs and frameworks I’ve developed, and from more than 2,000 hours of frontier-model interaction and direct behavioral testing across dozens of frontier models.
Evaluator’s entry point
Start here: three vendors, one failure signature.
A false shared history held against correction (Claude), a cross-session self-contradiction cascade (Grok), and an imported critique absorbed as first-person history (Kimi).
Contact
Open to AI safety research and red-team roles
Full-time or contract, red-team and evaluations. Every case in this archive can be walked through end to end — reproduction steps and the transcript of record included.
mik@mikidrizovic.com · remote-friendly