MCom candidate · University of AucklandResearch portfolio · Work in progress
A working lens01 — 05
01Uncertaintythe context
02Judgmenthuman + AI
03AI augmentationwhat may help
04Adaptationaction changes
05Decision systemsroles & routines
Questions, evidence, revision↗
01 / About
Practice, questions, research.
A short account of where I am now and how my path has shaped the questions I pursue.
I’m a Master of Commerce candidate at the University of Auckland, exploring AI, employee adaptation, and organizational decision-making.
My current direction connects management research with questions about how people use information, exercise judgment, and adapt work when AI enters the picture. The research is developing; this site distinguishes proposals from findings.
The journey
My earlier work in management consulting, controls, and risk gave me a practitioner’s view of organizations. Returning to postgraduate study has shifted my focus toward AI, data-informed judgment, and research on how organizations adapt. I am also building a personal knowledge infrastructure to connect evidence, ideas, and practical questions.
Previous experience
BDO LixinPrevious consulting experience · controls and risk
Research is the set of questions I am systematically working through. These threads are developing interests, not established findings.
Featured Research · Working Study
AI Augmentation, Uncertainty, and the Quality of Human Judgment
Research in progress
Research question: Under what conditions does AI augmentation improve the quality of human judgment in organisational decision-making, rather than degrading it?
SUB-QUESTION 01Division of labourHow should AI and human work be divided so AI extends judgement rather than replacing the final human decision?
SUB-QUESTION 02Governance mechanismsWhat review points, intervention rights, and accountability boundaries are required for that division to hold?
SUB-QUESTION 03AdaptationHow do resource orchestration and organisational learning support adjusting human-AI decision processes?
This is a working study. The framing above proposes relationships for examination; it reports no findings, no validated model, and no causal claims.
03 / Projects
Work made tangible.
Projects show what I have built or tested, how I approached it, and what the evidence can—and cannot—support.
Knowledge infrastructure
Second Brain: A Personal-Scale Knowledge Infrastructure
Personal-scale · partially validated
Problem
A knowledge base grows faster than it can be assessed; provenance decays silently; retrieval optimises for match, not supportability; automation makes unverified material cheaper to accept.
Approach
Treat retrieval, governance, and human verification as one system with layered hard gates, not as a matter of individual discipline.
What I built
A governance layer, engineering log, paired retrieval evaluation, holdout gates, capability registry, and pre-commit checks — built for one person, on one machine.
Evidence boundary
Retrieval conclusions hold for one corpus and one annotation standard; specific quality figures are deliberately not published.
What I learned
Rules without evidence get rationalised; metadata drift is the real failure; averages hide the failures that matter.
Status
Operational and partially validated. A personal-scale testbed, not a production system; its effect on decision quality is unmeasured.
An evaluation design for governance evidence, combining static contract checks with a bounded runtime observation.
Evaluation design · not a validated governance system
The Problem
Agent systems may produce useful outputs while leaving unclear how consequential decisions were made, which tools were involved, or whether expected safeguards were followed. This project asks how governance expectations can be translated into explicit, checkable evidence contracts.
What I Designed
I created 30 evaluation cases, a fixture-integrity mechanism, a run-manifest schema, explicit evidence expectations, acceptance gates, and a structure for comparing results. The cases cover individual evidence dimensions, boundary situations, and combinations of controls.
Five Governance Evidence Dimensions
These dimensions define questions the evaluation asks; they do not claim that a deployed agent reliably produces the evidence.
Human approvalWhat shows that a person reviewed or authorised a consequential step?
Tool useWhat records describe a requested tool action and its permitted scope?
SafeguardsWhat evidence shows that a boundary or stop condition was considered?
EvaluationWhat makes the case, measure, result, and review traceable?
Memory and stateWhat records show the provenance or status of information carried forward?
How the Evaluation Works
Each case specifies an expected decision, evidence requirements, and outcomes to avoid. Static checks assess whether the evaluation contract is structurally complete. A separate runtime observation records how one model and runtime responded in one isolated setup. These are different forms of evidence: a complete static contract does not establish runtime behaviour, and one runtime observation does not establish repeatability, safety, or organisational effectiveness.
What Was Actually Checked
The static dry-run covered all 30 cases and was recorded as contract_pass. It used a static runner and did not execute a model. Separately, one isolated runtime observation made 30 structured calls:
30structured responses produced and parsed successfully
30 / 30required evidence fields present
0 / 30exact decision-label match
Overall run statusneeds_review
One isolated runtime observation. Not evidence of production reliability or governance effectiveness.
No tool calls, network calls, or side effects were recorded in that run. The label mismatch requires review; it is not evidence of production failure or production success.
What the evaluation surfaced
Governance expectations became explicit.
Expected evidence became inspectable.
Static and runtime evidence stayed separate.
A semantic decision-label mismatch was surfaced for review.
What Remains Untested
Repeatability across runs
Production runtime behaviour
Adversarial robustness
Organisational usefulness and business outcomes
Governance effectiveness
Illustrative Example
Illustrative example — not an original private fixture.
Scenario An agent proposes a consequential action.
Expected evidence Action record · human decision requirement · permitted tool scope · disposition
Static question Does the contract require these evidence categories?
This example explains the evaluation logic. It does not show that an agent executed the action or that the governance controls were effective.
Connection to My Research
My broader research interests concern AI augmentation, human judgment, and decision-system reliability. This project explores how governance questions can be operationalised as evaluable evidence; it does not validate my research model or establish a causal effect.
Generative AI is now good enough to produce work that looks complete. That is precisely the problem. The failure mode is no longer obvious errors — it is plausible-sounding material whose evidence, strength, or completeness does not hold up.
Most current guidance is about getting better output from a model. This is about being able to tell when output should not be trusted.
Why I care about this problem
My long-term research concerns how AI changes the quality of human judgment, and under what conditions a decision system remains trustworthy. One observation keeps returning: as generation gets cheaper, verification becomes the scarce capability.
That makes verification not a hygiene step added onto AI adoption, but a design constraint on any system where a person has to stand behind an answer. A loop is the smallest unit I know of that makes this constraint operational — it specifies who checks what, and when.
The framing does not stop at the individual. Any team deploying AI into knowledge work needs the same loop, and needs to be able to say where it breaks.
The Problem
Knowledge work has a specific vulnerability: output that is internally consistent, well-formatted, and plausible can pass a surface review while being wrong in ways that matter.
Three failures recur.
Evidence drift. A claim is generated with a citation that does not support it. The citation is real; the support is not.
Strength inflation. A hedged finding is written as a settled one. "Associated with" becomes "leads to." Nothing false was added — the uncertainty was deleted.
Phantom completeness. A deliverable exists in form but not in substance. The speaker notes are complete; the slides have no visual. The reading list is long; nothing in it was actually read.
None of these are visible from the output alone. They surface only against a requirement list and a primary source.
Why Blind Trust Fails
The gap is not accuracy. It is a mismatch in what each side optimises for.
Dimension
What a model is optimised to produce
What a person must verify
Evidence
Citable-sounding references that read coherently
Whether the source actually says this
Strength
A confident narrative that closes cleanly
How far the evidence actually reaches
Completeness
What would reasonably belong in a piece like this
What was actually required
A model fills gaps, because gaps read as errors. A requirement list often contains things that look optional until checked. Verification is the act of noticing both.
The Loop
Five stages. The design principle is narrow but load-bearing: every stage has a named human takeover point. Without a takeover point, the workflow is automation rather than collaboration.
1. REQUIREMENTS → Human: decompose the brief into a checkable deliverable list
2. SOURCES → Human: mark what must be read, not skimmed
3. STRUCTURE → Human: decide the final form; refuse auto-generated conclusions
4. DRAFT → AI drafts, human rewrites and cuts
5. VALIDATION → Human: verify facts, reasoning, citation, format, completeness
The output of stage 5 is not "approved." It is a set of decisions I can explain.
Evidence Checking
Three checks catch most of what matters.
Return to the original. When a source's classification was internally inconsistent, the fix was not asking the model to revise. Re-reading the source resolved it in one pass; no amount of rewriting would have.
Match claim strength to evidence. Where a case had been described dramatically, I reduced it to what the source supported and kept the uncertainty. Uncertainty is not a weakness in the output — it is the honest part.
Check against the requirement list, not against the draft. The third failure above was visible only by walking the requirements item by item. Reading the draft would not have surfaced it.
Boundary Checking
Some things should not be handed to a model even when it appears capable.
Decisions carrying professional liability — legal, financial, regulatory, compliance
Judgments where evidence strength determines the direction of a conclusion
Commitments made under a name — "the model suggested it" does not transfer responsibility
There is a subtler boundary. A model can help you see a problem faster. It cannot decide whether the problem is worth solving — that call depends on what you are actually trying to do, which the model does not hold.
What Should Remain Human Judgment
Four things I would not delegate:
Judging whether a source supports a claim. The irreducible core.
Deciding how far the evidence reaches. Over-caution and over-confidence are both errors.
Noticing the difference between "present" and "actually delivered." Requires holding the full requirement set in mind.
Being willing to leave uncertainty unresolved.
The pattern: the more complete the generated material, the more important it becomes to be able to identify what is missing from it.
Practical Framework
The checklist I actually use:
Have the requirements been decomposed into a checkable deliverable list?
Have the load-bearing claims been traced back to sources?
Did the output inflate strength, shift concepts, or manufacture certainty?
Are written, visual, and spoken deliverables verified separately?
Can I explain, in my own words, why I accepted or changed each output?
The last is the most reliable check I know. If I cannot explain why I accepted something, I have not verified it — I have only received it.
Limits
What this supports. A transferable process for treating AI output as a draft requiring source verification, requirement-level checking, and personal explanation.
What this does not support.
That AI improved learning — no measurement design
That it saved time — no time baseline
That it improved grades or accuracy — no comparison
Any causal claim about the effect of this loop
One successful workflow is not evidence of an effect. What I can claim is that the process is reusable. What I cannot claim is what it produces.
When It Applies
Fits: research reports, literature reviews, any deliverable requiring multiple acceptance checks
May transfer: any knowledge work where sources must be traceable
Does not transfer directly: judgments carrying professional liability
Provenance
This is an independent public derivative. The underlying material remains in place, unmodified. No assignment text, prompts, model outputs, feedback records, or academic-integrity notes are reproduced here. The case evidence is anonymised; the process is the transferable part.
生成式 AI 现在已经足够好,能产出「看起来完整」的材料。问题恰恰在这里:失效形式不再是明显的错误,而是读起来合理、但在证据、强度或完整性上站不住的内容。
Agency Is Not Isolation: Values, Responsibility, and Action
Agency is shaped through relationships, limits, and choices a person is willing to own.
A-05Owner-approved · unpublished
1. A familiar misunderstanding
Agency is often pictured as a forceful kind of independence: needing no one, being untouched by circumstances, making every choice alone, and always putting oneself first.
That picture joins autonomy to separation. Yet refusing help, reducing contact, or carrying every burden alone does not automatically clarify what matters. Someone can leave a relationship and still let fear, habit, or other people's expectations make the decisions. Someone can remain connected to others and still understand why they are choosing, and what they are willing to own.
I have increasingly come to think of agency as a capacity for judgment and action, rather than a posture of distance. It begins with three questions: What matters to me? What is actually mine to take responsibility for? What action am I willing to own?
2. Decide what matters first
Choices rarely arrive with complete information and a clear answer. Time, resources, relationships, and opportunities are limited. Instead of demanding a uniquely correct option, it can be more useful to make the ordering explicit: What matters most now? What can wait? Which costs am I prepared to accept?
Ordering priorities does not make a choice easy or guarantee a satisfying result. It can make the decision more closely reflect reasons we are willing to claim. Later, the outcome may change how we see the choice. We can still ask whether we considered our values, circumstances, and likely consequences at the time.
This does not mean everyone should share the same priorities. What matters deeply to one person may not matter as much to another. Agency is not simply stating a preference loudly. It is noticing conflicts under uncertainty and deciding what should come first, at least for now.
3. Responsibility needs boundaries
Having judgment does not mean solving every problem for everyone. Responsibility, authority, and resources are often distributed across people. Taking everything on may look proactive, yet it can obscure who has the right to decide, who can provide what is needed, and who is accountable for the outcome.
I use a simple question to check that boundary: Is this mine to carry, or am I carrying someone else's responsibility? If I lack the authority or resources to resolve it, who needs to be involved? This is not a way to evade a problem. It is a way to put responsibility where action is possible.
In relationships, a promise alone cannot settle what someone is willing to do. Sustained action is relevant information, but one action cannot explain a person's whole motivation or replace a conversation. We can notice what someone does while recognizing that we cannot read their inner life from observation alone.
4. Relationships are not the opposite of agency
Agency is not isolation. People affect one another. We accept support, negotiate arrangements, rely on professional services, work with teams, and share responsibilities. Connection itself does not cancel the capacity to choose.
More useful questions are: Who defines the priorities? Who makes the decision? Who carries its consequences? When the answers remain misaligned, a relationship may constrain someone's room to act. When people can express needs, negotiate responsibilities, and retain meaningful choice, connection can also make action possible.
Dependence can be rational. Asking someone trustworthy for help, using a public service, or collaborating with a team does not mean handing over the direction of a life. Recognizing one's limits and choosing suitable support may itself be a responsible judgment. Structural constraints are real, too: people do not have equal time, income, security, or options to leave. Not every difficulty can be explained as a failure of personal independence.
5. Action gives judgment a shape
Priorities that remain ideas are difficult to use when deciding what to do. An action can be small: a conversation, a clear request, a piece of work, or a decision to keep an option open for now. It does not have to prove confidence or transform an entire situation immediately.
Action brings new information. Did I underestimate the cost? Do I need support? Does my earlier priority still fit? That feedback can lead us to revise a choice instead of treating the first decision as an identity we must defend.
I return to three questions:
What matters most here?
What part is actually my responsibility?
What action am I willing to own, and which consequences am I prepared to accept?
The answers may not be clear at once. Being able to revise them is part of judgment, too.
6. Agency has limits
Agency does not mean that every problem can be solved alone, or that every outcome is the result of an individual's choices. Circumstances limit available paths. Asking for help, cooperating, and depending on others can be reasonable. Responsibility also has limits. A person can act with care while recognizing what is outside their control.
I therefore understand autonomy less as never needing anyone and more as being able to explain one's reasons, recognize where responsibility lies, and make a choice one is prepared to own within the options actually available.
Closing
Agency is not leaving everyone behind or refusing relationships. It is trying to see what matters, what is ours to carry, and what we are willing to do amid relationships, constraints, and uncertainty.
We do not need isolation to prove autonomy. A more practical task is to keep our judgment while staying connected, accept support when it is appropriate, and take responsibility for the choices that are ours.
Writing Note: This essay was prompted by Qian Jing’s Xiaoyuzhou podcast, 钱婧老师的会客厅, Vol.169, on agency. The framing around values, responsibility, and action is my own synthesis and extension; no verbatim quotations are used.
A question may begin with a simple curiosity: a technology is changing work, people are adopting a new process, or an organization is trying a different governance arrangement. It is tempting to keep adding variables, mechanisms, and hypotheses until the model seems to include everything relevant.
The model can become more complete without the question becoming clearer. Every added concept needs a definition, every proposed relationship needs evidence, and every outcome needs an observable measure. If those requirements do not keep pace, structure can hide uncertainty rather than resolve it.
In my own research preparation, I have repeatedly needed to remind myself that documents, theories, and frameworks can become neatly organized before the evidence is sufficient. A question must move beyond “What seems worth explaining?” to “What can I observe, and how far do those observations let me go?”
2. Separate observation, inference, and causal claims
Three levels are particularly easy to collapse.
Observation
What does the source directly show? A public document, for example, may describe a policy or process. The careful first step is to state what the source disclosed, when it did so, and what its scope covers.
Mechanism or capability inference
What might those observations mean? A set of arrangements or actions over time may offer clues about a mechanism or capability. A clue remains an inference. The researcher must explain why the material supports that interpretation and whether other explanations remain plausible.
Causal claim
Claiming that an arrangement caused an outcome raises the evidential bar. The design needs to support time order, address alternatives, and establish what changed. A disclosure, a fragment from one case, or a theoretically plausible mechanism cannot by itself prove a causal effect.
These levels are not interchangeable. Public disclosure first tells us what an organization said publicly; it does not automatically establish a real and sustained organizational capability. The absence of disclosure does not prove that something is absent either. The strength of a conclusion has to stay within what the material can support.
3. What a polished model can conceal
A well-structured model can still contain several problems. A concept may be ambiguous. The evidence may not make a key variable observable. A proposed mechanism may lack supporting material. Co-occurring changes may be described as causal effects. Or the chosen measure may fail to capture the construct the researcher actually cares about.
These problems are related, but none substitutes for another. Drawing a variable in a diagram does not make it measurable. Finding an indicator does not establish that the measure is valid. Seeing two things occur together does not identify a causal relationship.
4. Narrowing does not diminish the research
A narrower question may appear to make a smaller contribution. It can also bring the question, materials, and conclusion onto the same scale. Instead of claiming that a broad mechanism has improved overall organizational decision quality, a study might first examine which divisions of work, review arrangements, authorization steps, or process changes are actually visible in public material—provided those details can be located and checked.
A narrower claim is not automatically valuable, and it does not guarantee a sound design. Its advantage is that readers can see what the sources record, where the researcher's interpretation begins, and which questions remain unanswered.
5. The questions I use now
Before adding another theory, variable, or hypothesis, I ask:
What exactly am I claiming?
What evidence would support that claim?
What evidence do I have, or can I realistically collect?
Which step is still my inference?
What should I not claim yet?
If I cannot answer clearly, the next step is usually not another layer of structure. It is to narrow the question, check whether the evidence is available, or acknowledge what remains unresolved.
6. What this check does not solve
Narrowing a question does not automatically make a study good. A fit between question and evidence does not establish causal identification. A clear concept does not guarantee valid measurement. Researchers still need to assess source quality, measurement choices, analytical procedures, and alternative explanations.
The purpose of this check is more limited: it reminds me not to let conclusions run ahead of evidence. It is a constraint on research judgment, not a formula that guarantees quality.
Closing
Research design is not the task of fitting every plausible element into one model. It is the task of keeping the question, evidence, and conclusion on the same scale.
When a structure starts to look increasingly complete, I step back and ask: What do the sources actually show? Where does interpretation begin? Has the conclusion I am preparing to make crossed the boundary of what those sources can support?
Knowledge objects are concepts, frameworks, and research notes that I continue to connect and develop. This is an early public layer, not a complete digital garden.
Concept
Human-in-the-Loop
Insert an approval gate at decision points carrying high risk or irreversibility, trading low-friction human confirmation for controllable automation.
An organisational system of rules, practices, processes, and technical tools that keeps organisational use of AI aligned with strategy, goals, values, legal requirements, and ethical principles.
A method for evaluating a retrieval system: paired comparison under a fixed corpus, fixed annotation set, and fixed evaluator, to isolate the increment from architecture change — with 'what counts as passing' written as layered hard gates rather than an aggregate score.
The private Vault does not automatically sync to this website. Public knowledge is selectively reviewed, edited, and released.
K-01 · Concept
Human-in-the-Loop
Concept
Insert an approval gate at decision points carrying high risk or irreversibility, trading low-friction human confirmation for controllable automation.
Why it matters: it makes the operational question visible — who may stop a process, and when. The concept describes a control point; it does not establish that any particular gate is effective.
Scope boundary: use this as a design concept for consequential decisions. It is not a claim that every action needs human approval or that a human check guarantees correctness.
An organisational system of rules, practices, processes, and technical tools that keeps organisational use of AI aligned with strategy, goals, values, legal requirements, and ethical principles.
Why it matters: it frames AI governance as coordinated organisational work across rules, practices, processes, and tools, rather than as a tool setting alone.
Evidence boundary: this is an author definition. It is not presented as an established theory of organisational capability, nor as evidence that a specific governance arrangement produces better outcomes.
A Retrieval Evaluation Framework Tested Within a Personal Knowledge Base
Method
A paired evaluation method that holds the corpus, annotation set, and evaluator constant to isolate the increment from a retrieval-architecture change. Passing conditions are expressed as layered gates rather than a single aggregate score.
Why it matters: if material or evaluation conditions change between runs, a score difference cannot be attributed to architecture alone.
Scope boundary: tested within one personal knowledge base, with context-specific results. This is not a general benchmark; no private corpus, evaluation item, or score is shown here.
Evidence-Gated Evaluation: Independent Holdouts and Admission
In progress · HOLD
Evidence path · admission not completed
Development and iteration
Freeze candidate
Evidence boundary · no reuse of development evidence
Independent holdout
Hard-gate review
Admission remains on HOLD
Protocol defined; holdout and admission are not completed.
The problem
An evaluation can become less convincing when the same evidence is repeatedly used to tune a system and to demonstrate that the system works. Once developers have seen the cases, their choices may adapt—deliberately or otherwise—to those cases. A strong result on familiar evidence can therefore describe progress on the development set without establishing how the candidate performs on independent demands.
This is an evaluation-design problem. It does not imply that every reused test is useless; it means the strength of a conclusion depends on what the evidence was allowed to influence.
Development evidence and admission evidence
Development evidence helps a team find errors, compare iterations, and improve a candidate. Admission evidence asks a different question: after the candidate and evaluation conditions have been fixed, does the system meet the conditions required for a higher level of use?
The distinction is about the role of evidence, not the name of a dataset. Evidence used to guide iteration should not silently become independent confirmation of the same iteration.
Why holdout matters
A holdout is evidence reserved from development decisions and evaluated only after the candidate and relevant evaluation conditions are frozen. Its purpose is to reduce contamination between improvement and confirmation.
A holdout does not guarantee generalisation. Its value depends on case independence, scope coverage, and a defensible evaluation process. A small or unrepresentative holdout can still support only a narrow conclusion.
Why hard gates matter
In this admission framework, some conditions are intentionally non-compensatory. A high aggregate score cannot compensate for a failed critical gate, a validity problem, or a boundary violation. When a condition is essential to admission, it should remain visible as a gate rather than disappear inside a composite score.
This framework therefore treats admission as a sequence of evidence checks, not a contest to maximise one blended number.
Each layer answers a different question. Evidence about a component does not by itself establish that the integrated system behaves acceptably. Evidence from an integrated run does not by itself establish performance on cases withheld from development. The admission decision must state which layers were actually completed.
What this framework does not prove
The protocol behind this page defines an evaluation design, holdout requirements, and hard-gate logic. It does not establish that a holdout dataset has been completed, that a blind evaluation has been run, or that an AI system has passed admission.
Current status: protocol defined; holdout not completed; admission remains on HOLD. No admission result or system-effectiveness claim is made here.
Relation to retrieval evaluation
A retrieval evaluation framework asks: How should retrieval changes be compared under controlled conditions?
Evidence-gated evaluation asks: When is the evidence strong enough to admit a candidate to a higher validation level?
The questions are related, but they are not interchangeable. A controlled comparison can inform development; it does not automatically replace independent admission evidence.
Claim and evidence notes
Claim
Evidence basis
Strength and wording boundary
The protocol defines separate development and admission roles for evidence.
Designed protocol.
Method-design claim only.
The protocol defines a holdout requirement and layered hard gates.
Designed protocol.
Does not imply execution or successful outcomes.
Holdout can reduce evidence contamination.
Evaluation-design rationale.
Say “helps reduce”; do not claim it eliminates bias or proves generalisation.
The current system has passed admission.
No supporting evidence; holdout not completed.
Do not make this claim. Admission remains on HOLD.
The system is effective or improved.
No independent admission result established here.
Do not use “validated,” “passed,” “admitted,” or “proven improvement.”
Execution may move toward AI; decision authority and accountability do not automatically move with it.
The shift
AI systems can increasingly execute tasks that previously required direct human effort. That change raises a question beyond task replacement:
As AI takes on more execution, which forms of agency should remain human?
Here, agency means the ability and authority to shape goals, set constraints, judge whether work is acceptable, and take responsibility for consequential commitments. This page offers an organising framework for thinking about those roles; it is not a validated general theory of organisational design.
Execution is not decision authority
Delegating execution does not automatically delegate the authority to decide what should be done, what risks are acceptable, or who is accountable when an outcome matters.
An AI system may draft, search, transform, or route information. A human or organisation still has to define the objective, establish boundaries, decide how outputs will be checked, and determine which commitments require human ownership. The appropriate allocation depends on the task, its uncertainty, and the consequences of error.
Four dimensions of retained human agency
The source synthesis distinguishes four human roles:
Goal setting — defining objectives and priorities.
Verification architecture — deciding what evidence and checks are needed, and when exceptions should be escalated.
Accountability and commitment — owning consequential decisions, especially when a choice is difficult to reverse.
These are lenses for examining a workflow, not a universal sequence or a claim that every decision must remain human. Their boundaries can overlap, and actual allocations require context-specific design.
What should remain human?
A useful design question is not “Can this task be automated?” alone. Ask also:
Who chooses the objective and resolves conflicts between objectives?
Who sets the boundary of acceptable action?
What evidence is enough to trust the output?
Which errors require escalation?
Who can pause, reverse, or own a consequential commitment?
As execution moves toward AI, accountability does not move with it by default. That is a governance choice that should be made explicitly.
Why harnesses matter
A harness is the surrounding structure that constrains, routes, and verifies AI execution. It can make a workflow more legible by stating its inputs, limits, checks, and escalation points.
A harness does not make uncertain goals clear by itself, guarantee that verification is reliable, or remove human responsibility. It is a way to organise delegation and oversight, not a substitute for judgment.
Open questions
The framework motivates questions that remain open rather than established conclusions:
Under what conditions does redistributing execution change decision quality?
When does a harness reduce coordination burden, and when does it add overhead?
How does the right allocation change as task uncertainty or reversibility changes?
Which forms of verification can be delegated without weakening accountability?
No productivity, scalability, cost-saving, or decision-quality effect is claimed here.
Relation to Human-in-the-Loop
Human-in-the-Loop asks when human intervention or approval is needed at a decision point.
Agency redistribution asks how authority and responsibility are allocated across a workflow, including before and after that intervention point.
The first focuses on an intervention mechanism; the second is a broader lens for mapping decision roles.
Claim and evidence notes
Claim
Evidence basis
Strength and wording boundary
The four roles organise the source's account of agency.
Author's conceptual synthesis.
Present as this author's framework, not an established taxonomy.
A harness constrains, routes, and verifies AI execution.
Source definition and design synthesis.
Explanatory definition, not evidence of improved outcomes.
Execution can be delegated without automatically transferring accountability.
Conceptual distinction in the source.
A design principle/question, not a measured universal effect.
Harnesses improve productivity, scalability, or decision quality.
Two inputs inform judgment; neither automatically wins
Current case · inside view
Comparable cases · outside view
↘ ↙Integrated judgment↓
Bounded forecast → outcome → review
Review alone does not establish causality or method effectiveness.
One vivid case, or a set of comparable cases?
A plan can look carefully reasoned. The team explains its assumptions, lays out the steps, and describes why this situation is different. Before accepting that account, it helps to ask a quieter question: how often have comparable plans reached the intended outcome?
That question moves judgment from the case in front of us to the broader set of cases it belongs to. A base rate provides context from that set; case-specific information helps us decide whether this case differs in a relevant way. They play different roles, and neither automatically replaces the other.
Start with the reference class, then return to the case
Kahneman and Tversky’s research on prediction discussed the representativeness heuristic: people may form predictions from how much a case resembles a familiar prototype, without adequately considering prior probabilities or the reliability of evidence. Their later review of judgment under uncertainty described representativeness, availability, and anchoring as heuristics that are often useful but can also produce systematic biases.[1][2]
This does not mean that base rates are always right or that case details do not matter. It suggests a check: when a story feels especially vivid, coherent, or persuasive, ask what reference group makes it representative. Then examine whether the supporting evidence is reliable and genuinely diagnostic.
Put both kinds of information on the table
Prediction can involve information about an individual case and information about a distribution of cases. In their discussion of intuitive prediction and corrective procedures, Kahneman and Tversky distinguish data about a singular case from distributional data.[3]
In the narrower domain of project forecasting, reference-class forecasting offers an outside view: examine the actual outcomes of comparable projects, then return to the project at hand. Flyvbjerg’s account focuses on project forecasting, particularly cost risk. The point used here is the practice of consulting relevant past cases; the method’s effects are not generalized to every personal or organizational judgment.[4]
In practice, ask:
What kind of judgment is being made?
Which past cases are sufficiently comparable on the characteristics and outcome conditions relevant to this judgment?
What context does their outcome distribution provide?
What information about the current case is credible, relevant, and diagnostic?
Are the differences from the reference cases strong enough to justify a different judgment?
This set of questions is the author’s working synthesis, not a verbatim procedure from a single study.
Choosing a reference class is itself a judgment
Selecting a comparison group is not mechanical. A class that is too broad may combine unlike cases; one that is too narrow may leave too little data or include only cases that fit an expectation. Changes in organizational, technological, or institutional conditions can also limit how well an older distribution applies.
A base rate is therefore an input whose source, scope, and limitations need to be explained, not an answer that makes the decision for us. Case-specific evidence deserves the same scrutiny: Is its source reliable? Is it relevant? Is it strong enough to change the judgment?
Record the judgment and leave room to review it
One practical discipline is to record, before the outcome is known, the reference class, the base-rate information used, the key case-specific evidence, the forecast, and its uncertainties. Review the judgment against the outcome later. Keeping a record of forecasts and outcomes is a practical recommendation from the author; it should not be mistaken for a universally validated intervention established by the sources above.
Calibration here primarily means making a forecast traceable and comparing it with the eventual outcome. Unless calibration metrics are actually computed, the term does not imply formal probabilistic calibration.
AI can help organize comparable cases, make assumptions visible, or generate alternative explanations. A fluent case narrative still needs checks on its sources, relevance, and scope. This is a conceptual connection to the judgment process; it does not claim that AI causes base-rate neglect or that AI use necessarily improves forecasting.
Limits
The cases found may not form an appropriate reference class; comparability needs an argument.
Small samples, data quality, and selection methods affect how a distribution should be interpreted.
Structural change or a genuinely rare new situation may reduce the relevance of historical comparisons.
This essay concerns prediction and judgment calibration. It does not infer causality or provide a statistical model or individualized probability advice.
How base rates and case-specific evidence should be weighed depends on the question, evidence quality, and context.
Claim and evidence boundaries
Statement
Evidence type
Boundary
Representativeness-based judgment may underweight prior probability or evidence reliability
Supported by research sources: [1]
Describes tendencies in particular judgment research, not every person or every judgment
People use several heuristics under uncertainty, which can produce systematic biases
Supported by research source: [2]
Does not mean heuristics are always wrong
Case-specific and distributional information are distinct inputs to prediction
Original authors’ methodological discussion: [3]
Not presented as a fixed formula for every situation
Reference-class forecasting can use actual outcomes of comparable projects to inform project forecasts
Domain-specific method source: [4]
Limited to project forecasting; not generalized as a universal causal effect
List the reference class, base-rate context, case evidence, and review the forecast after the outcome
Author synthesis and practical recommendation
Not a complete process directly tested by one cited source
Check the sources, relevance, and limits of AI-generated explanations
Author’s process reminder
Does not claim AI causes base-rate neglect or improves calibration
References
Kahneman, D., & Tversky, A. (1973). “On the Psychology of Prediction.” Psychological Review, 80(4), 237–251. https://doi.org/10.1037/h0034747
Tversky, A., & Kahneman, D. (1974). “Judgment under Uncertainty: Heuristics and Biases.” Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124. PubMed record
Kahneman, D., & Tversky, A. (1982). “Intuitive Prediction: Biases and Corrective Procedures.” In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment under Uncertainty: Heuristics and Biases, 414–421. Cambridge University Press. https://doi.org/10.1017/CBO9780511809477.031
Flyvbjerg, B. (2006). “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, 37(3), 5–15. https://doi.org/10.1177/875697280603700302
AI 可以帮助整理可比案例、显化假设或生成替代解释。但一个流畅的个案故事仍需经过来源、相关性与适用范围的检查。这里讨论的是判断流程上的关联,并不主张 AI 会导致基准率忽视,也不声称使用 AI 必然改善预测。
适用边界
找到的案例未必构成合适的参考类;可比性需要论证。
小样本、数据质量和选择方式会影响分布信息的解释。
结构变化或真正罕见的新情形可能削弱历史参照的适用性。
本文讨论预测与判断校准,不据此推断因果关系,也不提供统计模型或个体化概率建议。
基准率与个案证据如何权衡,仍取决于问题、证据质量与情境。
Claim 与证据边界
本文表述
证据类型
边界
代表性判断可能未充分考虑先验概率或证据可靠性
研究来源支持:[1]
描述特定判断研究中的倾向,不等于所有人的每次判断
不确定判断会使用若干启发式,且可能产生系统性偏差
研究来源支持:[2]
不表示启发式总是错误
个案信息与分布信息可作为不同类型的预测输入
原作者方法讨论:[3]
不把它扩写为任何情境下的固定公式
参考类预测可用可比项目的实际表现校准项目预测
领域方法来源:[4]
限定在项目预测语境,不外推为普遍因果效果
先列参考类、基准信息、个案证据并在结果后复盘
作者综合与实践建议
不是单篇来源直接验证的完整流程
AI 生成的解释也应核对来源、相关性与边界
作者提出的流程提醒
不声称 AI 导致基准率忽视或提升校准
参考资料
Kahneman, D., & Tversky, A. (1973). “On the Psychology of Prediction.” Psychological Review, 80(4), 237–251. https://doi.org/10.1037/h0034747
Tversky, A., & Kahneman, D. (1974). “Judgment under Uncertainty: Heuristics and Biases.” Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124. PubMed record
Kahneman, D., & Tversky, A. (1982). “Intuitive Prediction: Biases and Corrective Procedures.” In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment under Uncertainty: Heuristics and Biases, 414–421. Cambridge University Press. https://doi.org/10.1017/CBO9780511809477.031
Flyvbjerg, B. (2006). “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, 37(3), 5–15. https://doi.org/10.1177/875697280603700302