Experiment

AI Labs · Research synthesis case study

Complete

June 2026

AI Research

Trust

Without

Competence

An AI-assisted synthesis of 45 sources on what happens when people trust AI financial advice they can't independently evaluate — verified line by line, with every failure in the process left on the record.

Sources indexed

45/50

Themes survived

4 (1 high)

Quotes verified

8/8

Failures logged

13

The question

A three-part thesis, built to be broken

I went in with a claim I wanted to test against the literature, not confirm.

“The population least equipped to evaluate financial advice is adopting and trusting AI financial advice fastest; the tools may erode the very competence they claim to build; and overconfidence turns that into measurable harm.”

Three independently checkable claims: decoupling (trust outpaces literacy), erosion (AI use degrades competence), harm (overconfidence + AI use produces damage). The synthesis existed to find out where each one holds, where it breaks, and what argues against it.

The verdict

Four themes survived. None survived unqualified.

Each was checked against quoted source text and put through an adversarial second read before it counted.

Partial

Medium

Trust drives adoption more than literacy

Holds for chat-style AI advice — and inverts for robo-advisors, whose users skew more literate. Defensible only once the product surface is named.

Hanson & Ott 2026 · Klingbeil 2024 ⚠ general-AI · Eichler & Schwab 2024

Partial

Low

AI reliance correlates with de-skilling

What survives: passive delegation doesn't build competence. The stronger “erosion” framing rests on general-AI studies — flagged as inference for finance.

Gerlich 2025 ⚠ · Shen & Tamkin 2026 ⚠ · Eichler & Schwab 2024

Partial

Low–Med

Overconfidence × AI → harm

The harm is an interaction, not a main effect — AI willingness tracked with lower fraud loss on average; damage concentrated only in the overconfident subgroup.

Chawla et al. 2026 · N=3,689 of 25,539, nationally representative

Verified

High

Literacy calibrates the shape of trust

The strongest finding: literacy doesn't change how much people trust AI — it changes whether trust moves at all in response to evidence.

Han & Ko 2025 · Stradi & Verdickt 2025, N=3,000 · Romeo & Conti 2025, PRISMA

Adversarial check

What the corpus argued back

A dedicated counter-thesis prompt, run to surface confirmation bias — mine and the corpus's.

01

AI may improve outcomes, not erode them

— following LLM advice moved most people closer to life-cycle-optimal behavior (Choukhmane et al. 2026).

02

Adoption isn't driven by the least-equipped

— robo-advisor users skew more literate than average; the field calls this a paradox (Nourallah et al. 2026).

03

Overconfidence alone doesn't cause harm

— the main effect was null; harm appeared only as an interaction with AI willingness (Chawla et al. 2026).

The failure layer

The ledger, not hidden in the middle

A clean process is usually a process nobody looked at hard. Every break, logged as it happened.

12

Paywalled URLs failed bulk import

→ 10 of 12 recovered via manual PDF upload

4

DOI redirects indexed error pages as if they were papers

→ caught by checking indexed character counts, replaced

3

PMC pages returned reCAPTCHA walls and indexed as if successful

→ swapped to open-access versions

2

ResearchGate sources access-blocked

→ dropped, arm coverage held by remaining sources

2

No free copy locatable (B10, C9)

→ absent, judged acceptable, logged

1

An affiliation rendered as an author — “Alinia AI et al.” was institution #9 in the real author list

→ caught only by a post-run fulltext audit

Division of labor

What stayed human

The actual subject of this case study. The model clustered; the judgment about what the evidence means didn't move.

The AI-assisted layer did

Indexed 45 sources, source-grounded

Surfaced candidate themes & contradictions

Mapped evidence grades per theme

Pulled quotes for human verification

Ran the adversarial counter-thesis pass

Saved every raw output unedited

Gaurav did — and only Gaurav

Verified each theme against source text

Weighted a 117-person survey against a 25,539-person sample differently

Downgraded de-skilling to “too inference-heavy” at full strength

Kept the counter-evidence that complicated his own thesis

Wrote every interpretation

The payoff

Six insight statements, one flagged as a guess

Each is anchored to a source defensible in an interview. The sixth is openly interpretation — the difference is the point.

1

AI financial advice is associated with protection on average

— the harm concentrates in the people most confident they don't need it.

2

Financial literacy doesn't change how much people trust AI

— it changes whether that trust moves in response to evidence.

3

The least-equipped users aren't on the safest AI surfaces

— they're on the most frictionless ones.

4

AI can improve financial outcomes without building competence

— the two are separable, not the same thing.

5

The mechanism of harm is skipped verification

, not wild risk-taking.

6

Interpretation

The feature that would protect the most-harmed user is the one they'd rate worst

— so the market won't build it unprompted.

Lab learnings

What this run taught, and what it changes

The findings belong to the corpus. These belong to the method — each one changed how the next experiment gets run.

Tooling

Successful indexing is not correct content.

reCAPTCHA walls and DOI error pages both imported with a green check. Verify by indexed character count, never by import status.

Tooling

Retrieval is not deterministic across prompts.

The same theme drew partly different source sets in prompt 1 and prompt 3. Treat any single pass as one sample; ask twice and diff before trusting a source list.

Verification

The citation errors that matter survive every plausibility check.

An affiliation was rendered as a lead author and read as completely real. Author fields need checking against source fulltext — at scale, automated against Crossref.

Verification

Quote fidelity does not survive without a human pass.

Close paraphrase gets returned as exact quotation. All 8 quotes matched verbatim, but the check cost an hour — and skipping it would have meant publishing on trust.

Method

The adversarial prompt changed the conclusion, not just the confidence.

Running a dedicated counter-thesis query moved the target population from the ignorant to the miscalibrated-confident. Build refutation into the protocol, not the review.

Judgment

The gap between raw output and verified finding is the work.

Two of four themes were downgraded on verification. A synthesis that ratifies everything the model produced hasn't been checked — it's been transcribed.

Carries forward

Protective design can't be optional, can't be informational, and can't depend on explanations — the overconfident opt out of verification, override warnings, and only high-literacy users benefit from explanation-repair. That constraint set became the kill filter for the next experiment.

Not a systematic review, not peer-reviewed — a documented synthesis run, credible because the failure layer is visible and the weak findings are labelled weak.

Stack: NotebookLM (retrieval) · notebooklm-py (unofficial CLI) · Claude Code (execution) · human verification (the part that doesn't automate)

2023 — 2025

Grey Project Studio

I ran an independent design and dev studio, taking on more than ten clients across SaaS, fintech, and operational software, from the first discovery call to developer handoff. For three enterprise clients I built reusable component libraries in Figma and Token Studio, which cut down a lot of repeat design work between sprints. I also ran heuristic evaluations and usability tests on client MVPs, and the fixes we made raised onboarding completion by 10 to 15 percent. Most of the time I was the only one translating between what the client's business needed and what was technically possible, so scoping, stakeholder calls, and keeping delivery on track fell to me by default.

By The Numbers
10+
Clients Delivered
3
enterprise design systems
10 - 15 %
Onboarding lift
100 %
Client Ownership
2023 — 2025

Grey Project Studio

I ran an independent design and dev studio, taking on more than ten clients across SaaS, fintech, and operational software, from the first discovery call to developer handoff. For three enterprise clients I built reusable component libraries in Figma and Token Studio, which cut down a lot of repeat design work between sprints. I also ran heuristic evaluations and usability tests on client MVPs, and the fixes we made raised onboarding completion by 10 to 15 percent. Most of the time I was the only one translating between what the client's business needed and what was technically possible, so scoping, stakeholder calls, and keeping delivery on track fell to me by default.

By The Numbers
10+
Clients Delivered
3
enterprise design systems
10 - 15 %
Onboarding lift
100 %
Client Ownership
Design SystemsBrand IdentityAI WorkflowsInteraction DesignUser ResearchPrototypingUX DesignProduct StrategyDesign SystemsBrand IdentityAI WorkflowsInteraction DesignUser ResearchPrototypingUX DesignProduct Strategy

LET'S MAKE SOMETHING

LET'S MAKE
SOMETHING

LET'S MAKE SOMETHING

THAT MATTERS.

THAT
MATTERS.

THAT MATTERS.