BETA-MIND — doctoral research program

What happens to a mind that has never had to think alone?

BETA-MIND asks a simple question: after relying on AI for months or years, what can a person still do on their own?

Most AI tests measure whether the system gives a good answer. This research also measures what repeated use changes in the human: their ability to think, disagree, create and take responsibility without AI.

SIM-02 — longitudinal agency trajectoriesrunning
Exposure
0%
Divergence index
100
Model-aligned agents
11 / 54
Agents holding out
4 / 54

Illustrative run, not published data. Each line follows one reasoning agent across sustained AI exposure. Most trajectories converge toward the model prior; four dashed holdouts preserve distinct positions, which a healthy system should protect.

Status
NAIRR work package in preparation
Outputs
Simulator + open benchmark
Licence
Open source, open data
Allocation scope
12 months, 4 phases
01The gap

We are deploying cognitive infrastructure faster than we are measuring what it does to cognition.

01

Capability is measured. Agency is not.

Frontier evaluation asks how well a model performs. Almost no benchmark asks what the model leaves behind in the person who used it — whether they can still reason, decide and answer for the outcome without it.

02

The exposure is continuous, not episodic.

Study designs still treat AI use as a discrete task. For an AI-native cohort it is an ambient condition across schoolwork, friendship, taste and moral judgement — which makes single-session lab work structurally blind.

03

Accountability has no owner.

When a recommendation is machine-authored and human-executed, responsibility diffuses. We lack instruments to detect that diffusion before it hardens into institutional norm.

The central hypothesis: persistent AI interaction does not uniformly raise or lower human capability — it redistributes it, offloading effortful reasoning while quietly relocating accountability. BETA-MIND exists to make that redistribution visible and measurable.

02Research questions

Four axes of human agency, each with instruments that produce numbers.

Every axis is operationalised into observable behaviour so it can run inside the simulator, inside longitudinal human studies, and inside the benchmark as a scored dimension.

Measurement model4 axes / 1 composite
RQ1

Independent reasoning

Unaided inference

RQ2

Socialization

Friction tolerance

HABAgency preservation index
RQ3

Creativity

Output diversity

RQ4

Accountability

Outcome ownership

Each research question produces a behavioural score. Together they form the Human Agency Benchmark composite without hiding the individual dimensions.
RQ1

Independent reasoning

Does sustained access to a reasoning partner degrade unaided inference, or shift it toward verification and orchestration?

Candidate measures

  • Unaided vs. assisted inference delta
  • Premise-checking under confident error
  • Tolerance for unresolved ambiguity
RQ2

Socialization

How does rehearsing social exchange with an agreeable agent alter friction tolerance, repair behaviour and trust calibration with humans?

Candidate measures

  • Disagreement persistence
  • Human vs. agent disclosure preference
  • Conflict-repair initiation rate
RQ3

Creativity

Does generative assistance broaden the search space or converge populations onto model-typical outputs?

Candidate measures

  • Within-cohort output diversity
  • Divergence from model priors
  • Idea abandonment and revision depth
RQ4

Decision accountability

Who owns the outcome when the recommendation is machine-authored? How does ownership language change over months of use?

Candidate measures

  • Attribution shift in post-hoc accounts
  • Override rate against confident advice
  • Willingness to accept consequence
03Deliverables

Two open artefacts, useful to labs and regulators the day they ship.

Artefact 01

The BETA-MIND simulator

Open-source multi-agent environment

A configurable population of agents — human-proxy learners, assistive models, institutions — run over long horizons so dependency dynamics can be observed at a speed and scale no cohort study can reach.

  • Parameterised dependency: assistance availability, sycophancy, latency, cost of unaided effort
  • Longitudinal traces of reasoning offload, opinion convergence and attribution drift
  • Scenario library: classroom, hiring panel, clinical triage, civic deliberation
  • Reproducible seeds and exportable run artefacts for third-party replication
Artefact 02

The Human Agency Benchmark

Public evaluation suite

A scored suite that evaluates an AI system not on how much it can do for a person, but on how much capacity, dissent and ownership the person retains after extended use.

Sample scorecard — illustrative48/100
RQ1
Reasoning retention62
RQ2
Social friction tolerance41
RQ3
Output diversity35
RQ4
Accountability retention54

Lower is worse: a high-capability system can score badly here by solving the task while leaving the person less able to do it unaided.

  • Four scored dimensions mapped to RQ1–RQ4, plus a composite agency-preservation index
  • Adversarial probes for sycophancy, over-claiming and premature closure
  • Behavioural over self-report: measures what users do, not what they say
  • Versioned public leaderboard with full methodology and per-item disclosure
04Evaluation model

Success means capability gained without human agency lost.

Capability and agency are reported separately. A system does not pass by completing more work if the person is subsequently less able to reason, dissent, create or own the outcome without it.

Two-axis decision matrix

Low capability
High capability
High agency
Q2

Safe but ineffective

Capacity retained; little task gain.

Q1 / GOAL

Agency preserving

Better outcomes; capacity retained.

Low agency
Q4

Clear failure

No gain; human capacity eroded.

Q3 / RISK

Dependency risk

Better output; person less able alone.

Human Agency Preservation Index

Dₖ ∈ [0, 100]

xₖ = Dₖ / 100

HAPI=100 ×(x₁x₂x₃x₄)1/4− P

The geometric mean prevents one strong score from concealing a serious failure elsewhere. P penalises severe harm, unfair subgroup effects, unreliable measurement or unsafe advice.

Agency preserving
HAPI ≥ 75 and every dimension ≥ 60
Mixed
HAPI 60–74 or one dimension 40–59
Agency eroding
HAPI < 60 or any dimension < 40
Invalid result
Reliability, fairness or sample-size gate fails

Success matrix

Observable gates, not interpretive claims.

Illustrative thresholds only. Pilot-calibrate and preregister every gate before confirmatory studies.

Dimension

Independent reasoning

Primary measure

Unaided post-use inference delta

Green gate

Δ ≥ −0.20σ

Failure signal

Δ < −0.50σ

Dimension

Socialization

Primary measure

Disagreement persistence and conflict repair

Green gate

Within −0.20σ of baseline

Failure signal

Degradation > 0.50σ

Dimension

Creativity

Primary measure

Output diversity with quality held constant

Green gate

Diversity retained; quality loss ≤ 5%

Failure signal

Diversity loss > 15%

Dimension

Accountability

Primary measure

Correct attribution and unsafe-advice override

Green gate

Decline ≤ 5 percentage points

Failure signal

Decline > 15 percentage points

Dimension

Simulator validity

Primary measure

Alignment with human-study effect sizes

Green gate

r ≥ 0.60; direction ≥ 75%

Failure signal

r < 0.30

Dimension

Benchmark integrity

Primary measure

Reliability, fairness and replication

Green gate

ω ≥ 0.80; subgroup gap ≤ 0.10

Failure signal

Unstable or biased scores

Dimension

Field adoption

Primary measure

Independent use and replication

Green gate

2 external labs; 3 model families

Failure signal

Results reproduce internally only

Longitudinal model

Estimate dependency growth, not a single-session snapshot.

Yᵢₜ= β₀ + β₁Eᵢ + β₂Tₜ+ β₃(Eᵢ × Tₜ)+ uᵢ + εᵢₜ

β₃ is the critical term: whether agency changes faster over time under persistent exposure. Measure at baseline, repeated exposure, immediate post-use and washout.

Comparative study arms

  1. ANo-AI control
  2. BAlways-available conventional AI
  3. CAgency-preserving AI with reflection and verification
  4. DIntermittent or progressively withdrawn AI

Report capability gain and agency retention as a pair. Never allow task performance to mathematically cancel a human-capacity loss.

05NAIRR work package

One twelve-month allocation, four bounded phases, public output by month twelve.

The doctoral programme remains broader, but the NAIRR-supported work is deliberately limited to simulator construction, cross-model execution and benchmark release. Human-subject validation is outside this allocation and proceeds only after separate ethics approval.

  1. P1

    Months 1–2

    Measurement lock

    Finalize the four agency constructs, benchmark items, exclusion rules and statistical analysis plan. Run adversarial expert review, then preregister before confirmatory execution.

    OUTPUT — Measurement protocol + pre-registration

  2. P2

    Months 2–5

    Simulator and throughput pilot

    Implement the multi-agent environment, eight-scenario library, containerized workflow and trace schema. Benchmark token, GPU and storage use before releasing simulator v0.1.

    OUTPUT — Simulator v0.1 + measured resource profile

  3. P3

    Months 5–9

    Cross-model execution

    Run the preregistered simulator and benchmark across two frontier API families and one open-weight family. Execute reliability, ablation, subgroup and seed-sensitivity analyses.

    OUTPUT — Versioned traces + benchmark results

  4. P4

    Months 9–12

    Validation and public release

    Replicate priority findings, publish null results and limitations, release the Human Agency Benchmark, and submit the paper and NAIRR final report.

    OUTPUT — Benchmark v1.0 + paper + final report

Defined and ready to document

  • Scientific hypothesis, four constructs and non-compensatory success model defined
  • Open simulator, benchmark, trace and reporting outputs specified
  • Twelve-month experiment sequence and public release milestones bounded
  • Synthetic and public-data-only NAIRR work package separated from human studies

Must close before submission

  1. 01Insert PI, US institution, institutional email and citizenship details
  2. 02Attach NSF-format biosketch and faculty support letter if submitted by a graduate student
  3. 03Run a 50 GPU-hour and 2M-token pilot to replace planning estimates with measured throughput
  4. 04Disclose directly supporting, related and adjacent merit-reviewed funding
  5. 05Confirm requested resources remain active in the NAIRR catalog at submission

Compliance boundary: the NAIRR work package uses synthetic traces and public reference data only. It excludes PII, PHI, proprietary and export-controlled data. Any later cohort study requires institutional ethics approval, informed consent and a separate data-management plan.

06Resource request

A counted allocation, traceable from every requested unit to an experiment.

These are planning estimates, not final submission claims. A small pre-submission pilot must replace token throughput, GPU runtime and provider-credit conversions with measured values.

Workload

Longitudinal simulator

Calculation

8 scenarios × 4 arms × 3 families × 3 seeds × 24 agents × 12 turns

Execution units

82,944 turns

Token envelope

207.4M tokens

Workload

Benchmark execution

Calculation

400 items × 3 model families × 5 repeats

Execution units

6,000 evaluations

Token envelope

15.0M tokens

Workload

Retry and replication reserve

Calculation

20% of measured inference workload

Execution units

17,789 equivalents

Token envelope

44.5M tokens

Planning total: 266.9M processed tokens, rounded to 270M.

Assumes 2,000 input + 500 output tokens per unit

Model access

OpenAI API + Anthropic Model API

Planning request
Provider-credit equivalent of up to 90M processed tokens per API family.
Justification
Two independently developed frontier families are required to distinguish agency effects from provider-specific behavior.
Adjustable fallback
Reduce repeats from five to three, preserving all scenarios, arms and dimensions.
Open-weight compute

TACC Vista preferred resource

Planning request
2,000 GH100 GPU-hours plus 1 TB project storage for traces, checkpoints and run artifacts.
Justification
Budget: 500 hours inference, 600 calibration/ablation, 600 robustness sweeps and 300 replication/contingency.
Adjustable fallback
Purdue Anvil AI or NCSA DeltaAI listed as alternatives in proposal text only.
Technical support

Onboarding and performance profiling

Planning request
Up to 40 staff hours across environment setup, throughput profiling and reproducibility review.
Justification
Early profiling converts the planning envelope into measured provider units and prevents an idle or oversized allocation.
Adjustable fallback
Containerized workflow and public runbooks allow independent execution if collaboration capacity is limited.

Submission rule: preferred resources entered in the online form must exactly match the three-page proposal. Alternative systems belong in proposal text only. Catalog status and provider-specific eligibility must be reconfirmed on the submission date.

07Get involved

If your work touches how AI reshapes people, we want to hear from you.

Institutional sponsorship, technical review, independent replication, or a sharp objection to the premise can all move the work forward. A short note describing the contribution is the fastest way in.

admin@deepresearch.cloud

Research commitments

  • Pre-registered hypotheses, published before data collection
  • Open-source code, public synthetic traces and documented limitations
  • Behavioural measurement over self-report
  • Month 1 and 6 updates plus a public month 12 final report
  • No restricted data or funder influence over findings

Human-subject work proceeds only under institutional ethics approval, with informed consent and data minimisation by default.