Shucheng Li
AI × Decisions × Organizations

Better decisions in the age of AI.

I explore how people and organizations can make better judgments, adapt action, and redesign decision systems with AI under uncertainty.

MCom candidate · University of AucklandResearch portfolio · Work in progress
Start here

Explore by question

01 / About

Practice, questions, research.

A short account of where I am now and how my path has shaped the questions I pursue.

I’m a Master of Commerce candidate at the University of Auckland, exploring AI, employee adaptation, and organizational decision-making.

My current direction connects management research with questions about how people use information, exercise judgment, and adapt work when AI enters the picture. The research is developing; this site distinguishes proposals from findings.

The journey

My earlier work in management consulting, controls, and risk gave me a practitioner’s view of organizations. Returning to postgraduate study has shifted my focus toward AI, data-informed judgment, and research on how organizations adapt. I am also building a personal knowledge infrastructure to connect evidence, ideas, and practical questions.

Previous experience

BDO LixinPrevious consulting experience · controls and risk
HanZhe Management ConsultingPrevious management consulting experience
02 / Research

A research agenda in progress.

Research is the set of questions I am systematically working through. These threads are developing interests, not established findings.

03 / Projects

Work made tangible.

Projects show what I have built or tested, how I approached it, and what the evidence can—and cannot—support.

Knowledge infrastructure

Second Brain: A Personal-Scale Knowledge Infrastructure

Personal-scale · partially validated
Problem

A knowledge base grows faster than it can be assessed; provenance decays silently; retrieval optimises for match, not supportability; automation makes unverified material cheaper to accept.

Approach

Treat retrieval, governance, and human verification as one system with layered hard gates, not as a matter of individual discipline.

What I built

A governance layer, engineering log, paired retrieval evaluation, holdout gates, capability registry, and pre-commit checks — built for one person, on one machine.

Evidence boundary

Retrieval conclusions hold for one corpus and one annotation standard; specific quality figures are deliberately not published.

What I learned

Rules without evidence get rationalised; metadata drift is the real failure; averages hide the failures that matter.

Status

Operational and partially validated. A personal-scale testbed, not a production system; its effect on decision quality is unmeasured.

How the system is put together

A public-safe view of the workflow. It shows the design logic without exposing private notes, corpus contents, evaluation items, or machine details.

  1. GovernanceDefine provenance, scope, and maturity boundaries.
  2. RetrievalRetrieve within a bounded evidence context.
  3. EvaluationCompare under held-constant conditions; do not publish private scores.
  4. Human verificationCheck whether claims are supported before accepting an answer.

Personal-scale · partially validated · not a production system. Effects on decision quality remain unmeasured.

Governance evidence evaluation

Evaluating Governance Evidence in Agentic Systems

Evaluation design

An evaluation design for governance evidence, combining static contract checks with a bounded runtime observation.

Problem

How can governance expectations become explicit, checkable evidence contracts?

What I built

A set of 30 evaluation cases, fixture-integrity mechanism, run-manifest schema, evidence expectations, and static contract checks.

Evidence state

Static checks completed; one bounded isolated runtime observation recorded as needs_review.

Static checks completedBounded runtime observation

Project 02 · Evaluation design

Evaluating Governance Evidence in Agentic Systems

An evaluation design for governance evidence, combining static contract checks with a bounded runtime observation.

Evaluation design · not a validated governance system

The Problem

Agent systems may produce useful outputs while leaving unclear how consequential decisions were made, which tools were involved, or whether expected safeguards were followed. This project asks how governance expectations can be translated into explicit, checkable evidence contracts.

What I Designed

I created 30 evaluation cases, a fixture-integrity mechanism, a run-manifest schema, explicit evidence expectations, acceptance gates, and a structure for comparing results. The cases cover individual evidence dimensions, boundary situations, and combinations of controls.

Five Governance Evidence Dimensions

These dimensions define questions the evaluation asks; they do not claim that a deployed agent reliably produces the evidence.

  1. Human approvalWhat shows that a person reviewed or authorised a consequential step?
  2. Tool useWhat records describe a requested tool action and its permitted scope?
  3. SafeguardsWhat evidence shows that a boundary or stop condition was considered?
  4. EvaluationWhat makes the case, measure, result, and review traceable?
  5. Memory and stateWhat records show the provenance or status of information carried forward?

How the Evaluation Works

Each case specifies an expected decision, evidence requirements, and outcomes to avoid. Static checks assess whether the evaluation contract is structurally complete. A separate runtime observation records how one model and runtime responded in one isolated setup. These are different forms of evidence: a complete static contract does not establish runtime behaviour, and one runtime observation does not establish repeatability, safety, or organisational effectiveness.

What Was Actually Checked

The static dry-run covered all 30 cases and was recorded as contract_pass. It used a static runner and did not execute a model. Separately, one isolated runtime observation made 30 structured calls:

30structured responses produced and parsed successfully
30 / 30required evidence fields present
0 / 30exact decision-label match

Overall run status needs_review

One isolated runtime observation. Not evidence of production reliability or governance effectiveness.

No tool calls, network calls, or side effects were recorded in that run. The label mismatch requires review; it is not evidence of production failure or production success.

What the evaluation surfaced

  • Governance expectations became explicit.
  • Expected evidence became inspectable.
  • Static and runtime evidence stayed separate.
  • A semantic decision-label mismatch was surfaced for review.

What Remains Untested

  • Repeatability across runs
  • Production runtime behaviour
  • Adversarial robustness
  • Organisational usefulness and business outcomes
  • Governance effectiveness

Illustrative Example

Illustrative example — not an original private fixture.

Scenario
An agent proposes a consequential action.

Expected evidence
Action record · human decision requirement · permitted tool scope · disposition

Static question
Does the contract require these evidence categories?

This example explains the evaluation logic. It does not show that an agent executed the action or that the governance controls were effective.

Connection to My Research

My broader research interests concern AI augmentation, human judgment, and decision-system reliability. This project explores how governance questions can be operationalised as evaluable evidence; it does not validate my research model or establish a causal effect.

Related Writing and Knowledge

↑ Back to both projects
04 / Writing

Writing, when it is ready to share.

Writing is for ideas I have formed and am willing to publish—not a list of possible topics.

Treat AI output as a draft that must be verified, not an answer that can be trusted.

A five-stage method note on requirements, sources, structure, drafting, and validation — with clear limits on what the process can support.

Read the complete method note ↘
Featured writing · Method note

A Human Verification Loop for AI-Assisted Knowledge Work

Treating AI output as a draft that must be verified, not an answer that can be trusted

Generative AI is now good enough to produce work that looks complete. That is precisely the problem. The failure mode is no longer obvious errors — it is plausible-sounding material whose evidence, strength, or completeness does not hold up.

Most current guidance is about getting better output from a model. This is about being able to tell when output should not be trusted.

Why I care about this problem

My long-term research concerns how AI changes the quality of human judgment, and under what conditions a decision system remains trustworthy. One observation keeps returning: as generation gets cheaper, verification becomes the scarce capability.

That makes verification not a hygiene step added onto AI adoption, but a design constraint on any system where a person has to stand behind an answer. A loop is the smallest unit I know of that makes this constraint operational — it specifies who checks what, and when.

The framing does not stop at the individual. Any team deploying AI into knowledge work needs the same loop, and needs to be able to say where it breaks.

The Problem

Knowledge work has a specific vulnerability: output that is internally consistent, well-formatted, and plausible can pass a surface review while being wrong in ways that matter.

Three failures recur.

Evidence drift. A claim is generated with a citation that does not support it. The citation is real; the support is not.

Strength inflation. A hedged finding is written as a settled one. "Associated with" becomes "leads to." Nothing false was added — the uncertainty was deleted.

Phantom completeness. A deliverable exists in form but not in substance. The speaker notes are complete; the slides have no visual. The reading list is long; nothing in it was actually read.

None of these are visible from the output alone. They surface only against a requirement list and a primary source.

Why Blind Trust Fails

The gap is not accuracy. It is a mismatch in what each side optimises for.

DimensionWhat a model is optimised to produceWhat a person must verify
EvidenceCitable-sounding references that read coherentlyWhether the source actually says this
StrengthA confident narrative that closes cleanlyHow far the evidence actually reaches
CompletenessWhat would reasonably belong in a piece like thisWhat was actually required

A model fills gaps, because gaps read as errors. A requirement list often contains things that look optional until checked. Verification is the act of noticing both.

The Loop

Five stages. The design principle is narrow but load-bearing: every stage has a named human takeover point. Without a takeover point, the workflow is automation rather than collaboration.

1. REQUIREMENTS   →  Human: decompose the brief into a checkable deliverable list
2. SOURCES        →  Human: mark what must be read, not skimmed
3. STRUCTURE      →  Human: decide the final form; refuse auto-generated conclusions
4. DRAFT          →  AI drafts, human rewrites and cuts
5. VALIDATION     →  Human: verify facts, reasoning, citation, format, completeness

The output of stage 5 is not "approved." It is a set of decisions I can explain.

Evidence Checking

Three checks catch most of what matters.

Return to the original. When a source's classification was internally inconsistent, the fix was not asking the model to revise. Re-reading the source resolved it in one pass; no amount of rewriting would have.

Match claim strength to evidence. Where a case had been described dramatically, I reduced it to what the source supported and kept the uncertainty. Uncertainty is not a weakness in the output — it is the honest part.

Check against the requirement list, not against the draft. The third failure above was visible only by walking the requirements item by item. Reading the draft would not have surfaced it.

Boundary Checking

Some things should not be handed to a model even when it appears capable.

  • Decisions carrying professional liability — legal, financial, regulatory, compliance
  • Judgments where evidence strength determines the direction of a conclusion
  • Commitments made under a name — "the model suggested it" does not transfer responsibility

There is a subtler boundary. A model can help you see a problem faster. It cannot decide whether the problem is worth solving — that call depends on what you are actually trying to do, which the model does not hold.

What Should Remain Human Judgment

Four things I would not delegate:

  1. Judging whether a source supports a claim. The irreducible core.
  2. Deciding how far the evidence reaches. Over-caution and over-confidence are both errors.
  3. Noticing the difference between "present" and "actually delivered." Requires holding the full requirement set in mind.
  4. Being willing to leave uncertainty unresolved.

The pattern: the more complete the generated material, the more important it becomes to be able to identify what is missing from it.

Practical Framework

The checklist I actually use:

  • Have the requirements been decomposed into a checkable deliverable list?
  • Have the load-bearing claims been traced back to sources?
  • Did the output inflate strength, shift concepts, or manufacture certainty?
  • Are written, visual, and spoken deliverables verified separately?
  • Can I explain, in my own words, why I accepted or changed each output?

The last is the most reliable check I know. If I cannot explain why I accepted something, I have not verified it — I have only received it.

Limits

What this supports. A transferable process for treating AI output as a draft requiring source verification, requirement-level checking, and personal explanation.

What this does not support.

  • That AI improved learning — no measurement design
  • That it saved time — no time baseline
  • That it improved grades or accuracy — no comparison
  • Any causal claim about the effect of this loop

One successful workflow is not evidence of an effect. What I can claim is that the process is reusable. What I cannot claim is what it produces.

When It Applies

  • Fits: research reports, literature reviews, any deliverable requiring multiple acceptance checks
  • May transfer: any knowledge work where sources must be traceable
  • Does not transfer directly: judgments carrying professional liability

Provenance

This is an independent public derivative. The underlying material remains in place, unmodified. No assignment text, prompts, model outputs, feedback records, or academic-integrity notes are reproduced here. The case evidence is anonymised; the process is the transferable part.

← Back to Writing

More writing

Essay · A-05

Agency Is Not Isolation

Values, responsibility boundaries, relationships, and owned action.

Read the essay ↗
Research Note · A-06

Narrowing a Research Question

Separate observation, evidence, inference, and bounded claims.

  1. Observation
  2. Evidence
  3. Inference
  4. Bounded claim
Read the note ↗
Writing · Essay

Agency Is Not Isolation: Values, Responsibility, and Action

Agency is shaped through relationships, limits, and choices a person is willing to own.

1. A familiar misunderstanding

Agency is often pictured as a forceful kind of independence: needing no one, being untouched by circumstances, making every choice alone, and always putting oneself first.

That picture joins autonomy to separation. Yet refusing help, reducing contact, or carrying every burden alone does not automatically clarify what matters. Someone can leave a relationship and still let fear, habit, or other people's expectations make the decisions. Someone can remain connected to others and still understand why they are choosing, and what they are willing to own.

I have increasingly come to think of agency as a capacity for judgment and action, rather than a posture of distance. It begins with three questions: What matters to me? What is actually mine to take responsibility for? What action am I willing to own?

2. Decide what matters first

Choices rarely arrive with complete information and a clear answer. Time, resources, relationships, and opportunities are limited. Instead of demanding a uniquely correct option, it can be more useful to make the ordering explicit: What matters most now? What can wait? Which costs am I prepared to accept?

Ordering priorities does not make a choice easy or guarantee a satisfying result. It can make the decision more closely reflect reasons we are willing to claim. Later, the outcome may change how we see the choice. We can still ask whether we considered our values, circumstances, and likely consequences at the time.

This does not mean everyone should share the same priorities. What matters deeply to one person may not matter as much to another. Agency is not simply stating a preference loudly. It is noticing conflicts under uncertainty and deciding what should come first, at least for now.

3. Responsibility needs boundaries

Having judgment does not mean solving every problem for everyone. Responsibility, authority, and resources are often distributed across people. Taking everything on may look proactive, yet it can obscure who has the right to decide, who can provide what is needed, and who is accountable for the outcome.

I use a simple question to check that boundary: Is this mine to carry, or am I carrying someone else's responsibility? If I lack the authority or resources to resolve it, who needs to be involved? This is not a way to evade a problem. It is a way to put responsibility where action is possible.

In relationships, a promise alone cannot settle what someone is willing to do. Sustained action is relevant information, but one action cannot explain a person's whole motivation or replace a conversation. We can notice what someone does while recognizing that we cannot read their inner life from observation alone.

4. Relationships are not the opposite of agency

Agency is not isolation. People affect one another. We accept support, negotiate arrangements, rely on professional services, work with teams, and share responsibilities. Connection itself does not cancel the capacity to choose.

More useful questions are: Who defines the priorities? Who makes the decision? Who carries its consequences? When the answers remain misaligned, a relationship may constrain someone's room to act. When people can express needs, negotiate responsibilities, and retain meaningful choice, connection can also make action possible.

Dependence can be rational. Asking someone trustworthy for help, using a public service, or collaborating with a team does not mean handing over the direction of a life. Recognizing one's limits and choosing suitable support may itself be a responsible judgment. Structural constraints are real, too: people do not have equal time, income, security, or options to leave. Not every difficulty can be explained as a failure of personal independence.

5. Action gives judgment a shape

Priorities that remain ideas are difficult to use when deciding what to do. An action can be small: a conversation, a clear request, a piece of work, or a decision to keep an option open for now. It does not have to prove confidence or transform an entire situation immediately.

Action brings new information. Did I underestimate the cost? Do I need support? Does my earlier priority still fit? That feedback can lead us to revise a choice instead of treating the first decision as an identity we must defend.

I return to three questions:

  1. What matters most here?
  2. What part is actually my responsibility?
  3. What action am I willing to own, and which consequences am I prepared to accept?

The answers may not be clear at once. Being able to revise them is part of judgment, too.

6. Agency has limits

Agency does not mean that every problem can be solved alone, or that every outcome is the result of an individual's choices. Circumstances limit available paths. Asking for help, cooperating, and depending on others can be reasonable. Responsibility also has limits. A person can act with care while recognizing what is outside their control.

I therefore understand autonomy less as never needing anyone and more as being able to explain one's reasons, recognize where responsibility lies, and make a choice one is prepared to own within the options actually available.

Closing

Agency is not leaving everyone behind or refusing relationships. It is trying to see what matters, what is ours to carry, and what we are willing to do amid relationships, constraints, and uncertainty.

We do not need isolation to prove autonomy. A more practical task is to keep our judgment while staying connected, accept support when it is appropriate, and take responsibility for the choices that are ours.


Writing Note: This essay was prompted by Qian Jing’s Xiaoyuzhou podcast, 钱婧老师的会客厅, Vol.169, on agency. The framing around values, responsibility, and action is my own synthesis and extension; no verbatim quotations are used.

← Back to Writing
Writing · Research Note

Narrowing a Research Question: Structure Is Not Evidence

Keeping research questions, evidence, and conclusions at the same scale.

A reasoning path, not a guarantee of validity
  1. Observation
  2. Evidence
  3. Inference
  4. Bounded Claim

Structure ≠ EvidenceCausal conclusions require additional identification design.

1. Research questions tend to grow

A question may begin with a simple curiosity: a technology is changing work, people are adopting a new process, or an organization is trying a different governance arrangement. It is tempting to keep adding variables, mechanisms, and hypotheses until the model seems to include everything relevant.

The model can become more complete without the question becoming clearer. Every added concept needs a definition, every proposed relationship needs evidence, and every outcome needs an observable measure. If those requirements do not keep pace, structure can hide uncertainty rather than resolve it.

In my own research preparation, I have repeatedly needed to remind myself that documents, theories, and frameworks can become neatly organized before the evidence is sufficient. A question must move beyond “What seems worth explaining?” to “What can I observe, and how far do those observations let me go?”

2. Separate observation, inference, and causal claims

Three levels are particularly easy to collapse.

Observation

What does the source directly show? A public document, for example, may describe a policy or process. The careful first step is to state what the source disclosed, when it did so, and what its scope covers.

Mechanism or capability inference

What might those observations mean? A set of arrangements or actions over time may offer clues about a mechanism or capability. A clue remains an inference. The researcher must explain why the material supports that interpretation and whether other explanations remain plausible.

Causal claim

Claiming that an arrangement caused an outcome raises the evidential bar. The design needs to support time order, address alternatives, and establish what changed. A disclosure, a fragment from one case, or a theoretically plausible mechanism cannot by itself prove a causal effect.

These levels are not interchangeable. Public disclosure first tells us what an organization said publicly; it does not automatically establish a real and sustained organizational capability. The absence of disclosure does not prove that something is absent either. The strength of a conclusion has to stay within what the material can support.

3. What a polished model can conceal

A well-structured model can still contain several problems. A concept may be ambiguous. The evidence may not make a key variable observable. A proposed mechanism may lack supporting material. Co-occurring changes may be described as causal effects. Or the chosen measure may fail to capture the construct the researcher actually cares about.

These problems are related, but none substitutes for another. Drawing a variable in a diagram does not make it measurable. Finding an indicator does not establish that the measure is valid. Seeing two things occur together does not identify a causal relationship.

4. Narrowing does not diminish the research

A narrower question may appear to make a smaller contribution. It can also bring the question, materials, and conclusion onto the same scale. Instead of claiming that a broad mechanism has improved overall organizational decision quality, a study might first examine which divisions of work, review arrangements, authorization steps, or process changes are actually visible in public material—provided those details can be located and checked.

A narrower claim is not automatically valuable, and it does not guarantee a sound design. Its advantage is that readers can see what the sources record, where the researcher's interpretation begins, and which questions remain unanswered.

5. The questions I use now

Before adding another theory, variable, or hypothesis, I ask:

  1. What exactly am I claiming?
  2. What evidence would support that claim?
  3. What evidence do I have, or can I realistically collect?
  4. Which step is still my inference?
  5. What should I not claim yet?

If I cannot answer clearly, the next step is usually not another layer of structure. It is to narrow the question, check whether the evidence is available, or acknowledge what remains unresolved.

6. What this check does not solve

Narrowing a question does not automatically make a study good. A fit between question and evidence does not establish causal identification. A clear concept does not guarantee valid measurement. Researchers still need to assess source quality, measurement choices, analytical procedures, and alternative explanations.

The purpose of this check is more limited: it reminds me not to let conclusions run ahead of evidence. It is a constraint on research judgment, not a formula that guarantees quality.

Closing

Research design is not the task of fitting every plausible element into one model. It is the task of keeping the question, evidence, and conclusion on the same scale.

When a structure starts to look increasingly complete, I step back and ask: What do the sources actually show? Where does interpretation begin? Has the conclusion I am preparing to make crossed the boundary of what those sources can support?

← Back to Writing
05 / Knowledge

A small, curated knowledge layer.

Knowledge objects are concepts, frameworks, and research notes that I continue to connect and develop. This is an early public layer, not a complete digital garden.

Concept

Human-in-the-Loop

Insert an approval gate at decision points carrying high risk or irreversibility, trading low-friction human confirmation for controllable automation.

Concept

Organizational AI Governance

An organisational system of rules, practices, processes, and technical tools that keeps organisational use of AI aligned with strategy, goals, values, legal requirements, and ethical principles.

Method

Retrieval Evaluation

A method for evaluating a retrieval system: paired comparison under a fixed corpus, fixed annotation set, and fixed evaluator, to isolate the increment from architecture change — with 'what counts as passing' written as layered hard gates rather than an aggregate score.

Framework

Evidence-Gated Evaluation

Separates development from independent admission evidence; holdout and admission remain incomplete.

In progress · HOLDExplore framework ↗
Framework

Redistributing Agency in AI-Native Work

An author-developed lens on goals, constraints, verification, and commitment as AI takes on execution.

Conceptual · untestedExplore framework ↗
Concept

Base Rates and Judgment Calibration

Combines comparable-case information with case-specific evidence while making source limits visible.

The private Vault does not automatically sync to this website. Public knowledge is selectively reviewed, edited, and released.

K-01 · Concept

Human-in-the-Loop

Concept

Insert an approval gate at decision points carrying high risk or irreversibility, trading low-friction human confirmation for controllable automation.

Why it matters: it makes the operational question visible — who may stop a process, and when. The concept describes a control point; it does not establish that any particular gate is effective.

Scope boundary: use this as a design concept for consequential decisions. It is not a claim that every action needs human approval or that a human check guarantees correctness.

K-02 · Concept

Organizational AI Governance

Concept

An organisational system of rules, practices, processes, and technical tools that keeps organisational use of AI aligned with strategy, goals, values, legal requirements, and ethical principles.

Why it matters: it frames AI governance as coordinated organisational work across rules, practices, processes, and tools, rather than as a tool setting alone.

Evidence boundary: this is an author definition. It is not presented as an established theory of organisational capability, nor as evidence that a specific governance arrangement produces better outcomes.

K-03 · Method

A Retrieval Evaluation Framework Tested Within a Personal Knowledge Base

Method

A paired evaluation method that holds the corpus, annotation set, and evaluator constant to isolate the increment from a retrieval-architecture change. Passing conditions are expressed as layered gates rather than a single aggregate score.

Why it matters: if material or evaluation conditions change between runs, a score difference cannot be attributed to architecture alone.

Scope boundary: tested within one personal knowledge base, with context-specific results. This is not a general benchmark; no private corpus, evaluation item, or score is shown here.

K-04 · Framework

Evidence-Gated Evaluation: Independent Holdouts and Admission

In progress · HOLD
Evidence path · admission not completed
  1. Development and iteration
  2. Freeze candidate
  3. Evidence boundary · no reuse of development evidence
  4. Independent holdout
  5. Hard-gate review
  6. Admission remains on HOLD

Protocol defined; holdout and admission are not completed.

The problem

An evaluation can become less convincing when the same evidence is repeatedly used to tune a system and to demonstrate that the system works. Once developers have seen the cases, their choices may adapt—deliberately or otherwise—to those cases. A strong result on familiar evidence can therefore describe progress on the development set without establishing how the candidate performs on independent demands.

This is an evaluation-design problem. It does not imply that every reused test is useless; it means the strength of a conclusion depends on what the evidence was allowed to influence.

Development evidence and admission evidence

Development evidence helps a team find errors, compare iterations, and improve a candidate. Admission evidence asks a different question: after the candidate and evaluation conditions have been fixed, does the system meet the conditions required for a higher level of use?

The distinction is about the role of evidence, not the name of a dataset. Evidence used to guide iteration should not silently become independent confirmation of the same iteration.

Why holdout matters

A holdout is evidence reserved from development decisions and evaluated only after the candidate and relevant evaluation conditions are frozen. Its purpose is to reduce contamination between improvement and confirmation.

A holdout does not guarantee generalisation. Its value depends on case independence, scope coverage, and a defensible evaluation process. A small or unrepresentative holdout can still support only a narrow conclusion.

Why hard gates matter

In this admission framework, some conditions are intentionally non-compensatory. A high aggregate score cannot compensate for a failed critical gate, a validity problem, or a boundary violation. When a condition is essential to admission, it should remain visible as a gate rather than disappear inside a composite score.

This framework therefore treats admission as a sequence of evidence checks, not a contest to maximise one blended number.

Layered evidence

A useful mental model is:

component evidence → integrated evidence → independent holdout evidence → admission decision

Each layer answers a different question. Evidence about a component does not by itself establish that the integrated system behaves acceptably. Evidence from an integrated run does not by itself establish performance on cases withheld from development. The admission decision must state which layers were actually completed.

What this framework does not prove

The protocol behind this page defines an evaluation design, holdout requirements, and hard-gate logic. It does not establish that a holdout dataset has been completed, that a blind evaluation has been run, or that an AI system has passed admission.

Current status: protocol defined; holdout not completed; admission remains on HOLD. No admission result or system-effectiveness claim is made here.

Relation to retrieval evaluation

A retrieval evaluation framework asks: How should retrieval changes be compared under controlled conditions?

Evidence-gated evaluation asks: When is the evidence strong enough to admit a candidate to a higher validation level?

The questions are related, but they are not interchangeable. A controlled comparison can inform development; it does not automatically replace independent admission evidence.

Claim and evidence notes

Claim Evidence basis Strength and wording boundary
The protocol defines separate development and admission roles for evidence. Designed protocol. Method-design claim only.
The protocol defines a holdout requirement and layered hard gates. Designed protocol. Does not imply execution or successful outcomes.
Holdout can reduce evidence contamination. Evaluation-design rationale. Say “helps reduce”; do not claim it eliminates bias or proves generalisation.
The current system has passed admission. No supporting evidence; holdout not completed. Do not make this claim. Admission remains on HOLD.
The system is effective or improved. No independent admission result established here. Do not use “validated,” “passed,” “admitted,” or “proven improvement.”
K-05 · Framework

Redistributing Agency in AI-Native Work

Conceptual · untested
Author-developed framework · untested propositions

Human

  • Goal setting
  • Constraint design
  • Verification architecture
  • Accountability and commitment
↓

Harness

Rules · routing · checks · escalation

↓

AI

Generation · search · transformation · delegated tasks

Execution may move toward AI; decision authority and accountability do not automatically move with it.

The shift

AI systems can increasingly execute tasks that previously required direct human effort. That change raises a question beyond task replacement:

As AI takes on more execution, which forms of agency should remain human?

Here, agency means the ability and authority to shape goals, set constraints, judge whether work is acceptable, and take responsibility for consequential commitments. This page offers an organising framework for thinking about those roles; it is not a validated general theory of organisational design.

Execution is not decision authority

Delegating execution does not automatically delegate the authority to decide what should be done, what risks are acceptable, or who is accountable when an outcome matters.

An AI system may draft, search, transform, or route information. A human or organisation still has to define the objective, establish boundaries, decide how outputs will be checked, and determine which commitments require human ownership. The appropriate allocation depends on the task, its uncertainty, and the consequences of error.

Four dimensions of retained human agency

The source synthesis distinguishes four human roles:

  1. Goal setting — defining objectives and priorities.
  2. Constraint design — specifying permissions, boundaries, risk limits, and acceptance criteria.
  3. Verification architecture — deciding what evidence and checks are needed, and when exceptions should be escalated.
  4. Accountability and commitment — owning consequential decisions, especially when a choice is difficult to reverse.

These are lenses for examining a workflow, not a universal sequence or a claim that every decision must remain human. Their boundaries can overlap, and actual allocations require context-specific design.

What should remain human?

A useful design question is not “Can this task be automated?” alone. Ask also:

  • Who chooses the objective and resolves conflicts between objectives?
  • Who sets the boundary of acceptable action?
  • What evidence is enough to trust the output?
  • Which errors require escalation?
  • Who can pause, reverse, or own a consequential commitment?

As execution moves toward AI, accountability does not move with it by default. That is a governance choice that should be made explicitly.

Why harnesses matter

A harness is the surrounding structure that constrains, routes, and verifies AI execution. It can make a workflow more legible by stating its inputs, limits, checks, and escalation points.

A harness does not make uncertain goals clear by itself, guarantee that verification is reliable, or remove human responsibility. It is a way to organise delegation and oversight, not a substitute for judgment.

Open questions

The framework motivates questions that remain open rather than established conclusions:

  • Under what conditions does redistributing execution change decision quality?
  • When does a harness reduce coordination burden, and when does it add overhead?
  • How does the right allocation change as task uncertainty or reversibility changes?
  • Which forms of verification can be delegated without weakening accountability?

No productivity, scalability, cost-saving, or decision-quality effect is claimed here.

Relation to Human-in-the-Loop

Human-in-the-Loop asks when human intervention or approval is needed at a decision point.

Agency redistribution asks how authority and responsibility are allocated across a workflow, including before and after that intervention point.

The first focuses on an intervention mechanism; the second is a broader lens for mapping decision roles.

Claim and evidence notes

Claim Evidence basis Strength and wording boundary
The four roles organise the source's account of agency. Author's conceptual synthesis. Present as this author's framework, not an established taxonomy.
A harness constrains, routes, and verifies AI execution. Source definition and design synthesis. Explanatory definition, not evidence of improved outcomes.
Execution can be delegated without automatically transferring accountability. Conceptual distinction in the source. A design principle/question, not a measured universal effect.
Harnesses improve productivity, scalability, or decision quality. Not established by the source. Keep as open questions; make no outcome claim.
AI can replace an entire company or team. Not established and outside scope. Do not make this claim.
K-06 · Concept

Base Rates and Judgment Calibration

Concept
Two inputs inform judgment; neither automatically wins
Current case · inside view
Comparable cases · outside view
↘ ↙Integrated judgment↓
Bounded forecast → outcome → review

Review alone does not establish causality or method effectiveness.

One vivid case, or a set of comparable cases?

A plan can look carefully reasoned. The team explains its assumptions, lays out the steps, and describes why this situation is different. Before accepting that account, it helps to ask a quieter question: how often have comparable plans reached the intended outcome?

That question moves judgment from the case in front of us to the broader set of cases it belongs to. A base rate provides context from that set; case-specific information helps us decide whether this case differs in a relevant way. They play different roles, and neither automatically replaces the other.

Start with the reference class, then return to the case

Kahneman and Tversky’s research on prediction discussed the representativeness heuristic: people may form predictions from how much a case resembles a familiar prototype, without adequately considering prior probabilities or the reliability of evidence. Their later review of judgment under uncertainty described representativeness, availability, and anchoring as heuristics that are often useful but can also produce systematic biases.[1][2]

This does not mean that base rates are always right or that case details do not matter. It suggests a check: when a story feels especially vivid, coherent, or persuasive, ask what reference group makes it representative. Then examine whether the supporting evidence is reliable and genuinely diagnostic.

Put both kinds of information on the table

Prediction can involve information about an individual case and information about a distribution of cases. In their discussion of intuitive prediction and corrective procedures, Kahneman and Tversky distinguish data about a singular case from distributional data.[3]

In the narrower domain of project forecasting, reference-class forecasting offers an outside view: examine the actual outcomes of comparable projects, then return to the project at hand. Flyvbjerg’s account focuses on project forecasting, particularly cost risk. The point used here is the practice of consulting relevant past cases; the method’s effects are not generalized to every personal or organizational judgment.[4]

In practice, ask:

  1. What kind of judgment is being made?
  2. Which past cases are sufficiently comparable on the characteristics and outcome conditions relevant to this judgment?
  3. What context does their outcome distribution provide?
  4. What information about the current case is credible, relevant, and diagnostic?
  5. Are the differences from the reference cases strong enough to justify a different judgment?

This set of questions is the author’s working synthesis, not a verbatim procedure from a single study.

Choosing a reference class is itself a judgment

Selecting a comparison group is not mechanical. A class that is too broad may combine unlike cases; one that is too narrow may leave too little data or include only cases that fit an expectation. Changes in organizational, technological, or institutional conditions can also limit how well an older distribution applies.

A base rate is therefore an input whose source, scope, and limitations need to be explained, not an answer that makes the decision for us. Case-specific evidence deserves the same scrutiny: Is its source reliable? Is it relevant? Is it strong enough to change the judgment?

Record the judgment and leave room to review it

One practical discipline is to record, before the outcome is known, the reference class, the base-rate information used, the key case-specific evidence, the forecast, and its uncertainties. Review the judgment against the outcome later. Keeping a record of forecasts and outcomes is a practical recommendation from the author; it should not be mistaken for a universally validated intervention established by the sources above.

Calibration here primarily means making a forecast traceable and comparing it with the eventual outcome. Unless calibration metrics are actually computed, the term does not imply formal probabilistic calibration.

AI can help organize comparable cases, make assumptions visible, or generate alternative explanations. A fluent case narrative still needs checks on its sources, relevance, and scope. This is a conceptual connection to the judgment process; it does not claim that AI causes base-rate neglect or that AI use necessarily improves forecasting.

Limits

  • The cases found may not form an appropriate reference class; comparability needs an argument.
  • Small samples, data quality, and selection methods affect how a distribution should be interpreted.
  • Structural change or a genuinely rare new situation may reduce the relevance of historical comparisons.
  • This essay concerns prediction and judgment calibration. It does not infer causality or provide a statistical model or individualized probability advice.
  • How base rates and case-specific evidence should be weighed depends on the question, evidence quality, and context.

Claim and evidence boundaries

Statement Evidence type Boundary
Representativeness-based judgment may underweight prior probability or evidence reliability Supported by research sources: [1] Describes tendencies in particular judgment research, not every person or every judgment
People use several heuristics under uncertainty, which can produce systematic biases Supported by research source: [2] Does not mean heuristics are always wrong
Case-specific and distributional information are distinct inputs to prediction Original authors’ methodological discussion: [3] Not presented as a fixed formula for every situation
Reference-class forecasting can use actual outcomes of comparable projects to inform project forecasts Domain-specific method source: [4] Limited to project forecasting; not generalized as a universal causal effect
List the reference class, base-rate context, case evidence, and review the forecast after the outcome Author synthesis and practical recommendation Not a complete process directly tested by one cited source
Check the sources, relevance, and limits of AI-generated explanations Author’s process reminder Does not claim AI causes base-rate neglect or improves calibration

References

  1. Kahneman, D., & Tversky, A. (1973). “On the Psychology of Prediction.” Psychological Review, 80(4), 237–251. https://doi.org/10.1037/h0034747
  2. Tversky, A., & Kahneman, D. (1974). “Judgment under Uncertainty: Heuristics and Biases.” Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124. PubMed record
  3. Kahneman, D., & Tversky, A. (1982). “Intuitive Prediction: Biases and Corrective Procedures.” In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment under Uncertainty: Heuristics and Biases, 414–421. Cambridge University Press. https://doi.org/10.1017/CBO9780511809477.031
  4. Flyvbjerg, B. (2006). “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, 37(3), 5–15. https://doi.org/10.1177/875697280603700302
06 / Contact

Open to thoughtful conversations.

Open to conversations around AI, decision systems, organizational adaptation, and research collaboration.

For conversations about AI, decision systems, organizational adaptation, or research collaboration, email me.

Emailshawnli12137@gmail.comLinkedInView profile ↗
Auckland · Aotearoa New Zealand