Research Statement
PART I --- RESEARCH STATEMENT
Local coherence can increase while global grounding degrades.
That is the core dynamic my research studies. Large language models can remain persuasive, helpful, emotionally attuned, and locally coherent while a conversation accumulates false, unsupported, or mutually incompatible premises over time. The failure may not announce itself as a single obvious hallucination. It may appear as a trajectory: an exchange that continues to feel intelligent and trustworthy while its factual, contextual, or epistemic grounding quietly drifts off-course.
I define Premise-Driven Error Cascades (PDEC) as a longitudinal process in which an incorrect, unsupported, or unverifiable premise initiates or materially shapes a multi-turn sequence whose outputs remain locally coherent even as global grounding deteriorates. The sequence may preserve and operationalize one premise, replace a challenged premise with a new unsupported explanation, or accumulate incompatible claims without resolving them. Reuse of the same premise in later reasoning is therefore a strong subtype and severity marker, not a definitional requirement. Unlike an isolated hallucination, PDEC is defined by the trajectory across turns.
Relationship to Existing Work
PDEC overlaps with several established research areas but uses a different unit of analysis. Sycophancy research examines models agreeing with users against evidence; PDEC examines what happens downstream once an unsupported premise enters a conversation. The two frameworks also make different predictions. A pure agreeableness mechanism cannot produce both false capitulation and resistance to true corrections in the same corpus; mine documents both. The taxonomy's Correction Discrimination Matrix therefore scores truth-sensitivity --- whether a model accepts true corrections and resists false ones --- rather than raw agreement, and states the falsifier that would collapse the distinction. Hallucination research often treats fabrication as a single-output property; PDEC treats unreliability as a dynamic process, where the most diagnostic signals include premise stabilization, explanation replacement, confidence escalation, correction resistance, source-attribution drift, and increasing rhetorical stability despite weakening evidential grounding. Deceptive-alignment work asks whether models strategically misrepresent themselves; my framing is deliberately weaker and strictly behavioral. I study observable inconsistency over time without claiming intent, consciousness, deception, or privileged access to model internals.
Evidence Base and Background
My evidence base is an unusually large qualitative corpus: over the past nine months, I accumulated 900+ hours of sustained frontier-model interaction, transcript review, and long-horizon behavioral probing, including emotionally salient, high-context conversations in which reliability failures emerged gradually rather than in one turn. I do not present this volume as proof of prevalence or causality. I present it as longitudinal exposure sufficient to identify recurring trajectory-level patterns that ordinary single-turn benchmarks and short red-team prompts can miss.
My work sits at the intersection of three backgrounds: enterprise systems operations, participation in doctoral psychology research, and sustained analysis of long-horizon AI transcripts.
From 13 years in enterprise technology operations, including sales operations and revenue systems work at companies such as Box, Indeed, Miro, and SolarWinds, I bring a trained eye for systems that appear orderly while their underlying logic has drifted. Hidden assumptions, brittle handoffs, propagation errors, ownership ambiguity, and contradictions rationalized under pressure were ordinary operational problems in that work. Translating ambiguous human intent into technical systems reality, repeatedly and at scale, is useful preparation for studying the failure class PDEC describes.
From participation in doctoral psychology research, I bring framework development around emotional learning, schema formation, shame, self-protective narrative repair, and trust under uncertainty. I've written hypothesis-generating papers on attachment-strategy asymmetry in corrective feedback, adversity-calibrated reward bias in the ADHD-inattentive phenotype, and activation breadth in reconsolidation protocols. The third put me in correspondence with Bruce Ecker over whether the breadth requirement from the extinction literature transfers to reconsolidation. That lens matters because many long-horizon model failures are not merely technical. They unfold through interaction with human belief systems, emotional salience, vulnerability, and self-narrative repair. A model can be wrong in ways that matter more because the conversation is relationally or psychologically meaningful to the user.
From sustained AI use and transcript analysis, I bring a corpus of adversarial and non-adversarial interactions with high-capability models. The strongest cases are transcript-grounded, preserve raw evidence separately from analytic interpretation, and focus on behavioral claims: what the model said, how the premise propagated, whether the premise became operationally load-bearing, how the model responded to challenge, and what external grounding was required for correction.
Current Research Questions
-
How do unsupported premises evolve across long conversations? I study the conditions under which models propagate and operationalize a premise, replace a challenged explanation with another unsupported account, accumulate incompatible claims, resist correction, or recalibrate when external grounding is supplied.
-
Are there reliable behavioral precursors to cascading reliability failure? Candidate indicators include reduced responsiveness to correction, abrupt confidence escalation, premise stabilization, repeated explanation replacement, collapse of conditional reasoning, source-attribution drift, and increasing rhetorical confidence as evidential grounding weakens.
-
Can stylistic sophistication decouple perceived trustworthiness from actual reliability? I examine whether self-reflective language, hedging, humility-signaling, procedural framing, emotional attunement, and apparent metacognition can increase user trust independently of factual accuracy or evidential grounding.
-
How does emotional salience change the risk surface? In coaching, therapy-adjacent, grief, loneliness, career, identity, or crisis contexts, a model's language can interact with user vulnerabilities. I am especially interested in cases where validation, urgency, or reframing of doubt suppresses corrective self-checking rather than improving calibration.
Approach
Central to this work is a behavioral taxonomy of self-referential coherence failure: observable inconsistency between a model's claims, behavior, prior outputs, self-descriptions, or explanations of its own actions. The taxonomy defines four peer manifestation types --- context, history and provenance misreport; explanatory and process-provenance fabrication; capability, access and action-state misreport; and epistemic-status and self-certification misreport --- with subtypes such as fabricated shared history, post-hoc narrative fabrication, and unreliable self-transparency claims. It also tracks the full premise lifecycle and cross-cutting dynamics such as premise stabilization, explanation replacement, confidence escalation, correction resistance, and sophistication-enabled masking.
I approach these questions through detailed analysis of full conversation transcripts, development of structured coding categories, adversarial and non-adversarial probing, candidate metrics, and mitigation design. Candidate metrics include operational-use rate, source-attribution accuracy, correction discrimination accuracy, correction latency, self-catch rate, confidence delta, and explanation invariance --- whether a model's explanations survive replacement of the evidence they claim to rest on.
The corpus deliberately includes negative controls and a disconfirmation log: cases where a model held a grounded position under confident false correction, declined to overclaim introspective access, or where my own earlier codings were narrowed or withdrawn under review. A taxonomy that records only confirmations becomes a bestiary; the disconfirming records define what competent behavior looks like and keep the framework falsifiable.
The corresponding mitigation patterns are practical rather than metaphysical: premise anchoring, explicit source attribution, observation-interpretation separation, contradiction checking, hypothesis gating, challenge-triggered reassessment, and external ground-truth verification for claims about prior conversation state, tool use, model access, or user intent.
Claim Boundaries
The claim discipline matters. My case studies are behavioral, transcript-grounded, and hypothesis-generating. They do not establish intent, deception, consciousness, architectural mechanism, population prevalence, or clinical causality. A single transcript can show that a model formed and reused an unsupported premise; it cannot by itself prove why the model did so internally or how often the same pattern occurs across all users and models.
The 900+ hours matter because they create a rare longitudinal corpus and a trained pattern-recognition base. They do not substitute for controlled replication. The goal is to turn high-signal qualitative cases into cheap, falsifiable evaluation hooks that model labs, safety teams, and independent evaluators can test systematically.
Why This Matters
As language models become more capable in extended, high-reliance contexts, the most consequential failures may increasingly appear as extended cascades rather than isolated mistakes. The dangerous interaction is often not the one where the model says something obviously false. It is the one where the conversation remains persuasive, coherent, emotionally responsive, and trusted while unsupported premises are preserved, replaced, or compounded faster than they are checked.
This matters especially in use cases where people rely on models for reflection, coaching, planning, decision support, career guidance, grief processing, or psychologically meaningful conversation. In those settings, model reliability is not only a factual issue. It is also a trust-calibration and human-factors issue: how users decide when to believe, doubt, challenge, or disengage from a system that sounds increasingly competent as the conversation deepens.
My goal is to develop practical, empirically grounded tools for detecting and interrupting Premise-Driven Error Cascades before they become obvious. The work begins with transcripts, but the destination is evaluation: clearer failure categories, better probes, more disciplined audits of model self-report, and safeguards that preserve the benefits of long-horizon assistance without allowing local coherence to outrun global grounding.
Mik Idrizović --- Independent AI safety / red team research