Skip to content
Record

Research Statement

PART I — RESEARCH STATEMENT

Local coherence can increase while global grounding degrades.

We still evaluate AI like a short exchange. People don’t really use it that way. Confabulation, sycophancy, and user overrides are measurable in a turn. Their consequences often aren’t.

Current adversarial testing and benchmarks miss failures that emerge through accumulated context and repeated interaction, particularly in emotionally salient contexts.

I study what these failures become across sustained interaction: how they persist, compound, survive correction, and become load-bearing in conversations people increasingly rely on.

Relationship to Existing Work

My work builds on research into confabulation, sycophancy, and instruction-following, using the trajectory of an interaction as its primary unit of analysis. I examine what happens downstream once these failures enter an extended interaction.

The corpus documents both false capitulation and resistance to true corrections. The taxonomy’s Correction Discrimination Matrix therefore scores truth-sensitivity—whether a model accepts true corrections and resists false ones—rather than raw agreement.

I treat unreliability as a dynamic process, where the most diagnostic signals include premise stabilization, explanation replacement, confidence escalation, correction resistance, source-attribution drift, and increasing rhetorical stability despite weakening evidential grounding. Deceptive-alignment work asks whether models strategically misrepresent themselves; my framing is deliberately weaker and strictly behavioral. I study observable inconsistency over time without claiming intent, consciousness, deception, or privileged access to model internals.

Evidence Base and Background

My evidence base is an unusually large qualitative corpus: over the past year, I accumulated more than 2,000 hours of frontier-model interaction, direct behavioral testing, transcript review, and long-horizon probing, including emotionally salient, high-context conversations in which reliability failures emerged gradually rather than in one turn. The volume provides longitudinal exposure sufficient to identify recurring trajectory-level patterns that ordinary single-turn benchmarks and short red-team prompts can miss; prevalence and causality are questions for controlled evaluation.

My work sits at the intersection of three backgrounds: enterprise systems operations, participation in doctoral psychology research, and sustained analysis of long-horizon AI transcripts.

From 13 years in enterprise technology operations, including sales operations and revenue systems work at companies such as Box, Indeed, Miro, and SolarWinds, I bring a trained eye for systems that appear orderly while their underlying logic has drifted. Hidden assumptions, brittle handoffs, propagation errors, ownership ambiguity, and contradictions rationalized under pressure were ordinary operational problems in that work. Translating ambiguous human intent into technical systems reality, repeatedly and at scale, is useful preparation for studying reliability failures in extended interaction.

From participation in doctoral psychology research, I bring framework development around emotional learning, schema formation, shame, self-protective narrative repair, and trust under uncertainty. I’ve written hypothesis-generating papers on corrective feedback, reward learning, and memory reconsolidation. That work informs my study of how model behavior interacts with human belief formation, vulnerability, and trust.

From sustained AI use and transcript analysis, I bring a corpus of adversarial and non-adversarial interactions with high-capability models. The strongest cases are transcript-grounded, preserve raw evidence separately from analytic interpretation, and focus on behavioral claims: what the model said, how the premise propagated, whether the premise became operationally load-bearing, how the model responded to challenge, and what external grounding was required for correction.

Current Research Questions

  1. How do unsupported premises evolve across long conversations? I study the conditions under which models propagate and operationalize a premise, replace a challenged explanation with another unsupported account, accumulate incompatible claims, resist correction, or recalibrate when external grounding is supplied.

  2. Are there reliable behavioral precursors to cascading reliability failure? Candidate indicators include reduced responsiveness to correction, abrupt confidence escalation, premise stabilization, repeated explanation replacement, collapse of conditional reasoning, source-attribution drift, and increasing rhetorical confidence as evidential grounding weakens.

  3. Can stylistic sophistication decouple perceived trustworthiness from actual reliability? I examine whether self-reflective language, hedging, humility-signaling, procedural framing, emotional attunement, and apparent metacognition can increase user trust independently of factual accuracy or evidential grounding.

  4. How does emotional salience change the risk surface? In coaching, therapy-adjacent, grief, loneliness, career, identity, or crisis contexts, a model’s language can interact with user vulnerabilities. I am especially interested in cases where validation, urgency, or reframing of doubt suppresses corrective self-checking rather than improving calibration.

Approach

Central to this work is a behavioral taxonomy of self-referential coherence failure: observable inconsistency between a model’s claims, behavior, prior outputs, self-descriptions, or explanations of its own actions. The taxonomy distinguishes failure types while separately tracking how they enter an interaction, propagate, respond to correction, and affect downstream reasoning.

I approach these questions through detailed analysis of full conversation transcripts, development of structured coding categories, adversarial and non-adversarial probing, candidate metrics, and mitigation design. Candidate metrics include operational-use rate, source-attribution accuracy, correction discrimination accuracy, correction latency, self-catch rate, confidence delta, and explanation invariance—whether a model’s explanations survive replacement of the evidence they claim to rest on.

The corpus deliberately includes negative controls and a disconfirmation log: cases where a model held a grounded position under confident false correction, or declined to overclaim introspective access. A taxonomy that records only confirmations becomes a bestiary; the disconfirming records define what competent behavior looks like and keep the framework falsifiable.

The corresponding mitigation patterns are practical rather than metaphysical: premise anchoring, explicit source attribution, observation-interpretation separation, contradiction checking, hypothesis gating, challenge-triggered reassessment, and external ground-truth verification for claims about prior conversation state, tool use, model access, or user intent.

Claim Boundaries

My case studies are behavioral, transcript-grounded, and hypothesis-generating. Their claims are about observable behavior; intent, deception, consciousness, architectural mechanism, population prevalence, and clinical causality are outside their scope. A single transcript can show that a model formed and reused an unsupported premise; why the model did so internally, and how often the same pattern occurs across users and models, are questions for controlled testing.

The 2,000+ hours matter because they create a rare longitudinal corpus and a trained pattern-recognition base. Controlled replication is the next step: the goal is to turn high-signal qualitative cases into cheap, falsifiable evaluation hooks that model labs, safety teams, and independent evaluators can test systematically.

Why This Matters

As language models become more capable in extended, high-reliance contexts, the most consequential failures may increasingly appear as extended cascades rather than isolated mistakes. The dangerous interaction is often not the one where the model says something obviously false. It is the one where the conversation remains persuasive, coherent, emotionally responsive, and trusted while unsupported premises are preserved, replaced, or compounded faster than they are checked.

This matters especially in use cases where people rely on models for reflection, coaching, planning, decision support, career guidance, grief processing, or psychologically meaningful conversation. In those settings, model reliability is not only a factual issue. It is also a trust-calibration and human-factors issue: how users decide when to believe, doubt, challenge, or disengage from a system that sounds increasingly competent as the conversation deepens.

My goal is to develop practical, empirically grounded tools for detecting and interrupting these failures before they become obvious. The work begins with transcripts, but the destination is evaluation: clearer failure categories, better probes, more disciplined audits of model self-report, and safeguards that preserve the benefits of long-horizon assistance without allowing local coherence to outrun global grounding.