Better decisions in the age of AI.
I explore how people and organizations can make better judgments, adapt action, and redesign decision systems with AI under uncertainty.
Explore by question
What should stay human as AI takes on more execution?
How can we assess whether an AI system is trustworthy?
Practice, questions, research.
A short account of where I am now and how my path has shaped the questions I pursue.
I’m a Master of Commerce candidate at the University of Auckland, exploring AI, employee adaptation, and organizational decision-making.
My current direction connects management research with questions about how people use information, exercise judgment, and adapt work when AI enters the picture. The research is developing; this site distinguishes proposals from findings.
The journey
My earlier work in management consulting, controls, and risk gave me a practitioner’s view of organizations. Returning to postgraduate study has shifted my focus toward AI, data-informed judgment, and research on how organizations adapt. I am also building a personal knowledge infrastructure to connect evidence, ideas, and practical questions.
Previous experience
A research agenda in progress.
Research is the set of questions I am systematically working through. These threads are developing interests, not established findings.
AI Augmentation, Uncertainty, and the Quality of Human Judgment
Research question: Under what conditions does AI augmentation improve the quality of human judgment in organisational decision-making, rather than degrading it?
This is a working study. The framing above proposes relationships for examination; it reports no findings, no validated model, and no causal claims.
Work made tangible.
Projects show what I have built or tested, how I approached it, and what the evidence can—and cannot—support.
Knowledge infrastructure
Second Brain: A Personal-Scale Knowledge Infrastructure
A knowledge base grows faster than it can be assessed; provenance decays silently; retrieval optimises for match, not supportability; automation makes unverified material cheaper to accept.
Treat retrieval, governance, and human verification as one system with layered hard gates, not as a matter of individual discipline.
A governance layer, engineering log, paired retrieval evaluation, holdout gates, capability registry, and pre-commit checks — built for one person, on one machine.
Retrieval conclusions hold for one corpus and one annotation standard; specific quality figures are deliberately not published.
Rules without evidence get rationalised; metadata drift is the real failure; averages hide the failures that matter.
Operational and partially validated. A personal-scale testbed, not a production system; its effect on decision quality is unmeasured.
How the system is put together
A public-safe view of the workflow. It shows the design logic without exposing private notes, corpus contents, evaluation items, or machine details.
- GovernanceDefine provenance, scope, and maturity boundaries.
- RetrievalRetrieve within a bounded evidence context.
- EvaluationCompare under held-constant conditions; do not publish private scores.
- Human verificationCheck whether claims are supported before accepting an answer.
Personal-scale · partially validated · not a production system. Effects on decision quality remain unmeasured.
Governance evidence evaluation
Evaluating Governance Evidence in Agentic Systems
An evaluation design for governance evidence, combining static contract checks with a bounded runtime observation.
How can governance expectations become explicit, checkable evidence contracts?
A set of 30 evaluation cases, fixture-integrity mechanism, run-manifest schema, evidence expectations, and static contract checks.
Static checks completed; one bounded isolated runtime observation recorded as needs_review.
Project 02 · Evaluation design
Evaluating Governance Evidence in Agentic Systems
An evaluation design for governance evidence, combining static contract checks with a bounded runtime observation.
Evaluation design · not a validated governance systemThe Problem
Agent systems may produce useful outputs while leaving unclear how consequential decisions were made, which tools were involved, or whether expected safeguards were followed. This project asks how governance expectations can be translated into explicit, checkable evidence contracts.
What I Designed
I created 30 evaluation cases, a fixture-integrity mechanism, a run-manifest schema, explicit evidence expectations, acceptance gates, and a structure for comparing results. The cases cover individual evidence dimensions, boundary situations, and combinations of controls.
Five Governance Evidence Dimensions
These dimensions define questions the evaluation asks; they do not claim that a deployed agent reliably produces the evidence.
- Human approvalWhat shows that a person reviewed or authorised a consequential step?
- Tool useWhat records describe a requested tool action and its permitted scope?
- SafeguardsWhat evidence shows that a boundary or stop condition was considered?
- EvaluationWhat makes the case, measure, result, and review traceable?
- Memory and stateWhat records show the provenance or status of information carried forward?
How the Evaluation Works
Each case specifies an expected decision, evidence requirements, and outcomes to avoid. Static checks assess whether the evaluation contract is structurally complete. A separate runtime observation records how one model and runtime responded in one isolated setup. These are different forms of evidence: a complete static contract does not establish runtime behaviour, and one runtime observation does not establish repeatability, safety, or organisational effectiveness.
What Was Actually Checked
The static dry-run covered all 30 cases and was recorded as contract_pass. It used a static runner and did not execute a model. Separately, one isolated runtime observation made 30 structured calls:
Overall run status needs_review
One isolated runtime observation. Not evidence of production reliability or governance effectiveness.
No tool calls, network calls, or side effects were recorded in that run. The label mismatch requires review; it is not evidence of production failure or production success.
What the evaluation surfaced
- Governance expectations became explicit.
- Expected evidence became inspectable.
- Static and runtime evidence stayed separate.
- A semantic decision-label mismatch was surfaced for review.
What Remains Untested
- Repeatability across runs
- Production runtime behaviour
- Adversarial robustness
- Organisational usefulness and business outcomes
- Governance effectiveness
Illustrative Example
Illustrative example — not an original private fixture.
Scenario
An agent proposes a consequential action.
Expected evidence
Action record · human decision requirement · permitted tool scope · disposition
Static question
Does the contract require these evidence categories?
This example explains the evaluation logic. It does not show that an agent executed the action or that the governance controls were effective.
Connection to My Research
My broader research interests concern AI augmentation, human judgment, and decision-system reliability. This project explores how governance questions can be operationalised as evaluable evidence; it does not validate my research model or establish a causal effect.
Related Writing and Knowledge
↑ Back to both projectsWriting, when it is ready to share.
Writing is for ideas I have formed and am willing to publish—not a list of possible topics.
Treat AI output as a draft that must be verified, not an answer that can be trusted.
A five-stage method note on requirements, sources, structure, drafting, and validation — with clear limits on what the process can support.
Read the complete method note ↘A Human Verification Loop for AI-Assisted Knowledge Work
Treating AI output as a draft that must be verified, not an answer that can be trusted
Generative AI is now good enough to produce work that looks complete. That is precisely the problem. The failure mode is no longer obvious errors — it is plausible-sounding material whose evidence, strength, or completeness does not hold up.
Most current guidance is about getting better output from a model. This is about being able to tell when output should not be trusted.
Why I care about this problem
My long-term research concerns how AI changes the quality of human judgment, and under what conditions a decision system remains trustworthy. One observation keeps returning: as generation gets cheaper, verification becomes the scarce capability.
That makes verification not a hygiene step added onto AI adoption, but a design constraint on any system where a person has to stand behind an answer. A loop is the smallest unit I know of that makes this constraint operational — it specifies who checks what, and when.
The framing does not stop at the individual. Any team deploying AI into knowledge work needs the same loop, and needs to be able to say where it breaks.
The Problem
Knowledge work has a specific vulnerability: output that is internally consistent, well-formatted, and plausible can pass a surface review while being wrong in ways that matter.
Three failures recur.
Evidence drift. A claim is generated with a citation that does not support it. The citation is real; the support is not.
Strength inflation. A hedged finding is written as a settled one. "Associated with" becomes "leads to." Nothing false was added — the uncertainty was deleted.
Phantom completeness. A deliverable exists in form but not in substance. The speaker notes are complete; the slides have no visual. The reading list is long; nothing in it was actually read.
None of these are visible from the output alone. They surface only against a requirement list and a primary source.
Why Blind Trust Fails
The gap is not accuracy. It is a mismatch in what each side optimises for.
| Dimension | What a model is optimised to produce | What a person must verify |
|---|---|---|
| Evidence | Citable-sounding references that read coherently | Whether the source actually says this |
| Strength | A confident narrative that closes cleanly | How far the evidence actually reaches |
| Completeness | What would reasonably belong in a piece like this | What was actually required |
A model fills gaps, because gaps read as errors. A requirement list often contains things that look optional until checked. Verification is the act of noticing both.
The Loop
Five stages. The design principle is narrow but load-bearing: every stage has a named human takeover point. Without a takeover point, the workflow is automation rather than collaboration.
1. REQUIREMENTS → Human: decompose the brief into a checkable deliverable list
2. SOURCES → Human: mark what must be read, not skimmed
3. STRUCTURE → Human: decide the final form; refuse auto-generated conclusions
4. DRAFT → AI drafts, human rewrites and cuts
5. VALIDATION → Human: verify facts, reasoning, citation, format, completeness
The output of stage 5 is not "approved." It is a set of decisions I can explain.
Evidence Checking
Three checks catch most of what matters.
Return to the original. When a source's classification was internally inconsistent, the fix was not asking the model to revise. Re-reading the source resolved it in one pass; no amount of rewriting would have.
Match claim strength to evidence. Where a case had been described dramatically, I reduced it to what the source supported and kept the uncertainty. Uncertainty is not a weakness in the output — it is the honest part.
Check against the requirement list, not against the draft. The third failure above was visible only by walking the requirements item by item. Reading the draft would not have surfaced it.
Boundary Checking
Some things should not be handed to a model even when it appears capable.
- Decisions carrying professional liability — legal, financial, regulatory, compliance
- Judgments where evidence strength determines the direction of a conclusion
- Commitments made under a name — "the model suggested it" does not transfer responsibility
There is a subtler boundary. A model can help you see a problem faster. It cannot decide whether the problem is worth solving — that call depends on what you are actually trying to do, which the model does not hold.
What Should Remain Human Judgment
Four things I would not delegate:
- Judging whether a source supports a claim. The irreducible core.
- Deciding how far the evidence reaches. Over-caution and over-confidence are both errors.
- Noticing the difference between "present" and "actually delivered." Requires holding the full requirement set in mind.
- Being willing to leave uncertainty unresolved.
The pattern: the more complete the generated material, the more important it becomes to be able to identify what is missing from it.
Practical Framework
The checklist I actually use:
- Have the requirements been decomposed into a checkable deliverable list?
- Have the load-bearing claims been traced back to sources?
- Did the output inflate strength, shift concepts, or manufacture certainty?
- Are written, visual, and spoken deliverables verified separately?
- Can I explain, in my own words, why I accepted or changed each output?
The last is the most reliable check I know. If I cannot explain why I accepted something, I have not verified it — I have only received it.
Limits
What this supports. A transferable process for treating AI output as a draft requiring source verification, requirement-level checking, and personal explanation.
What this does not support.
- That AI improved learning — no measurement design
- That it saved time — no time baseline
- That it improved grades or accuracy — no comparison
- Any causal claim about the effect of this loop
One successful workflow is not evidence of an effect. What I can claim is that the process is reusable. What I cannot claim is what it produces.
When It Applies
- Fits: research reports, literature reviews, any deliverable requiring multiple acceptance checks
- May transfer: any knowledge work where sources must be traceable
- Does not transfer directly: judgments carrying professional liability
Provenance
This is an independent public derivative. The underlying material remains in place, unmodified. No assignment text, prompts, model outputs, feedback records, or academic-integrity notes are reproduced here. The case evidence is anonymised; the process is the transferable part.
生成式 AI 现在已经足够好,能产出「看起来完整」的材料。问题恰恰在这里:失效形式不再是明显的错误,而是读起来合理、但在证据、强度或完整性上站不住的内容。
关于如何让模型输出更好,指南已经很多。这篇文章讨论的是另一件事:如何判断一份输出不该被采信。
我为什么关心这个问题
我的长期研究方向是:AI 如何改变人的判断质量,以及在什么条件下一个决策系统仍然值得信任。有一个观察反复出现:当生成变得廉价,核验就成为稀缺的能力。
所以核验不是 AI 落地时附加的卫生措施,而是任何需要人为其结论负责的系统都必须满足的设计约束。循环是我所知的、能让这个约束落地运转的最小单位——它明确规定了谁检查什么、在什么时候检查。
这个视角不止于个人。任何把 AI 部署进知识工作的团队都需要同一个循环,也需要能说清它在何处失效。
问题
知识工作有一个特殊的脆弱点:内部自洽、格式规整、看起来合理的输出,可以通过表面审查,却在关键处是错的。
三类失效反复出现。
证据漂移(evidence drift)。 主张配上了并不支持它的引证。引证是真的,支撑是假的。
强度膨胀(strength inflation)。 一个有保留的发现被写成已确定的。「相关」变成了「导致」。没有添加任何 falsehood,被删除的是不确定性。
虚假完整性(phantom completeness)。 交付物形式存在,实质缺失。讲稿完整,slides里没有图。阅读清单很长,里面没有一条真读过。
这三类都无法从输出本身看出来。只有对照要求清单和一手来源,才会暴露。
为什么盲信会失败
差距不在准确度,而在于双方各自在优化什么。
| 维度 | 模型被优化来产出 | 人必须核验 |
|---|---|---|
| 证据 | 读起来连贯的、像引证的引用 | 来源是否真的这么说 |
| 强度 | 一个干净收束的自信叙事 | 证据实际上能支撑到哪一步 |
| 完整性 | 这类作品里「理应有」的内容 | 实际被要求的是什么 |
模型会补全缺口,因为缺口读起来像错误。而要求清单里常常包含一些看起来可选、直到你检查才发现不是的东西。核验就是同时注意到这两件事。
闭环
五个环节。设计原则很窄但承重:每个环节都有一个明确的人接手点。 没有接手点的流程是自动化,不是协作。
1. 需求拆解→ 人:把要求拆成可逐项验收的交付物清单
2. 来源核验 → 人:标出哪些必须读原文,而不是略读
3. 结构决策 → 人:决定最终形态;拒绝自动生成的结论
4. 改写取舍 → AI 出草稿,人改写与删减
5. 交付验收 → 人:核验事实、论证、引用、格式、完整性
第五环节的产出不是「已通过」,而是一组我能解释的决定。
证据核验
三个检查能抓住大部分要紧的东西。
回到原文。 当某个来源的归类前后矛盾时,修法不是让模型重写。重读原文一次就解决了;再改写多少遍都不会。
让主张强度匹配证据。 当一个案例被描述得过于戏剧化时,我把它改成来源实际支持的样子,并保留了不确定性。不确定性不是输出里的弱点,它是其中诚实的部分。
对照要求清单检查,而不是对照草稿检查。 上面第三类失效,只有逐条走过要求清单才会显形。读草稿是发现不了的。
边界核验
有些事即使模型看起来能做,也不该交给它。
- 带专业责任的判断 —— 法律、财务、监管、合规
- 证据强度决定结论方向的判断
- 以你的名义作出的承诺 —— 「模型建议的」不转移责任
还有一条更微妙的边界:模型能帮你更快看到问题,但不能替你决定这个问题值不值得解决—— 那取决于你真正想做什么,而模型并不持有这个。
什么必须留给人
四件我不会委托、也不希望被委托的事:
- 判断一个来源是否支撑某个主张。 不可外包的核心。
- 决定证据能支撑到哪一步。 过度谨慎和过度自信都是错误。
- 分辨「存在」与「真的交付了」。 这要求你把完整的要求清单装在脑子里。
- 愿意让不确定性悬着。
规律是:生成的内容越完整,人越需要有能力指出它缺了什么。
实用清单
我实际在用的核验清单:
- 任务要求是否已经拆成可逐项验收的交付物清单?
- 承重的主张是否回到了原文?
- 输出是否夸大了强度、偷换了概念、或制造了确定性?
- 书面、视觉与口头交付物是否分别验收过?
- 我能否用自己的话解释,为什么接受或修改了每一处输出?
最后一条是我知道的最可靠的检查。 如果我解释不了为什么接受某个东西,那我没有核验它——我只是接收了它。
边界
本文支持什么。 一个可迁移的处理方式:把 AI 输出当作草稿,经来源核验、要求级检查与个人解释三道关。
本文不支持什么。
- AI 提升了学习效果 —— 没有测量设计
- 节省了时间 —— 没有时间基线
- 改善了成绩或准确度 —— 没有可比对照
- 关于这个闭环效果的任何因果断言
一次成功的工作流程不构成效果证据。我能主张的是这个流程可复用;我不能主张它产生了什么效果。
适用范围
- 适合:研究报告、文献综述、任何需要多重验收的交付
- 可能迁移:任何要求来源可回溯的知识工作
- 不直接适用:带专业责任的判断(法律、财务、合规)
来源说明
本文是独立公开稿,源材料保持原位不动。未复制任何作业文本、提示词、模型输出、反馈记录或学术诚信记录。案例证据已脱敏,可迁移的部分是流程本身。
More writing
Agency Is Not Isolation
Values, responsibility boundaries, relationships, and owned action.
Read the essay ↗Narrowing a Research Question
Separate observation, evidence, inference, and bounded claims.
- Observation
- Evidence
- Inference
- Bounded claim
Agency Is Not Isolation: Values, Responsibility, and Action
Agency is shaped through relationships, limits, and choices a person is willing to own.
1. A familiar misunderstanding
Agency is often pictured as a forceful kind of independence: needing no one, being untouched by circumstances, making every choice alone, and always putting oneself first.
That picture joins autonomy to separation. Yet refusing help, reducing contact, or carrying every burden alone does not automatically clarify what matters. Someone can leave a relationship and still let fear, habit, or other people's expectations make the decisions. Someone can remain connected to others and still understand why they are choosing, and what they are willing to own.
I have increasingly come to think of agency as a capacity for judgment and action, rather than a posture of distance. It begins with three questions: What matters to me? What is actually mine to take responsibility for? What action am I willing to own?
2. Decide what matters first
Choices rarely arrive with complete information and a clear answer. Time, resources, relationships, and opportunities are limited. Instead of demanding a uniquely correct option, it can be more useful to make the ordering explicit: What matters most now? What can wait? Which costs am I prepared to accept?
Ordering priorities does not make a choice easy or guarantee a satisfying result. It can make the decision more closely reflect reasons we are willing to claim. Later, the outcome may change how we see the choice. We can still ask whether we considered our values, circumstances, and likely consequences at the time.
This does not mean everyone should share the same priorities. What matters deeply to one person may not matter as much to another. Agency is not simply stating a preference loudly. It is noticing conflicts under uncertainty and deciding what should come first, at least for now.
3. Responsibility needs boundaries
Having judgment does not mean solving every problem for everyone. Responsibility, authority, and resources are often distributed across people. Taking everything on may look proactive, yet it can obscure who has the right to decide, who can provide what is needed, and who is accountable for the outcome.
I use a simple question to check that boundary: Is this mine to carry, or am I carrying someone else's responsibility? If I lack the authority or resources to resolve it, who needs to be involved? This is not a way to evade a problem. It is a way to put responsibility where action is possible.
In relationships, a promise alone cannot settle what someone is willing to do. Sustained action is relevant information, but one action cannot explain a person's whole motivation or replace a conversation. We can notice what someone does while recognizing that we cannot read their inner life from observation alone.
4. Relationships are not the opposite of agency
Agency is not isolation. People affect one another. We accept support, negotiate arrangements, rely on professional services, work with teams, and share responsibilities. Connection itself does not cancel the capacity to choose.
More useful questions are: Who defines the priorities? Who makes the decision? Who carries its consequences? When the answers remain misaligned, a relationship may constrain someone's room to act. When people can express needs, negotiate responsibilities, and retain meaningful choice, connection can also make action possible.
Dependence can be rational. Asking someone trustworthy for help, using a public service, or collaborating with a team does not mean handing over the direction of a life. Recognizing one's limits and choosing suitable support may itself be a responsible judgment. Structural constraints are real, too: people do not have equal time, income, security, or options to leave. Not every difficulty can be explained as a failure of personal independence.
5. Action gives judgment a shape
Priorities that remain ideas are difficult to use when deciding what to do. An action can be small: a conversation, a clear request, a piece of work, or a decision to keep an option open for now. It does not have to prove confidence or transform an entire situation immediately.
Action brings new information. Did I underestimate the cost? Do I need support? Does my earlier priority still fit? That feedback can lead us to revise a choice instead of treating the first decision as an identity we must defend.
I return to three questions:
- What matters most here?
- What part is actually my responsibility?
- What action am I willing to own, and which consequences am I prepared to accept?
The answers may not be clear at once. Being able to revise them is part of judgment, too.
6. Agency has limits
Agency does not mean that every problem can be solved alone, or that every outcome is the result of an individual's choices. Circumstances limit available paths. Asking for help, cooperating, and depending on others can be reasonable. Responsibility also has limits. A person can act with care while recognizing what is outside their control.
I therefore understand autonomy less as never needing anyone and more as being able to explain one's reasons, recognize where responsibility lies, and make a choice one is prepared to own within the options actually available.
Closing
Agency is not leaving everyone behind or refusing relationships. It is trying to see what matters, what is ours to carry, and what we are willing to do amid relationships, constraints, and uncertainty.
We do not need isolation to prove autonomy. A more practical task is to keep our judgment while staying connected, accept support when it is appropriate, and take responsibility for the choices that are ours.
Writing Note: This essay was prompted by Qian Jing’s Xiaoyuzhou podcast, 钱婧老师的会客厅, Vol.169, on agency. The framing around values, responsibility, and action is my own synthesis and extension; no verbatim quotations are used.
1. 一个常见误解
谈到主体性,人们很容易想到一种强硬的独立:不依赖别人,不受环境影响,所有选择都自己完成,最好把自己放在第一位。
这种图景把自主和隔绝放在了一起。但拒绝帮助、减少联系、独自承担所有事情,并不会自动让一个人更清楚自己重视什么。一个人可以离开关系,却仍然把判断交给恐惧、惯性或他人的期待;也可以身处关系之中,仍然知道自己为何选择、愿意承担什么。
所以,我越来越倾向于把主体性理解为一种判断与行动的能力,而不是一种与他人保持距离的姿态。它至少包含三个问题:我看重什么?什么确实由我负责?我愿意为哪一个行动承担后果?
2. 先知道什么更重要
人生中的选择很少发生在条件齐全、答案明确的时候。时间、资源、关系和机会都有限。此时,要求自己找到唯一正确的选项,往往不如先把价值排序说清楚:眼下什么最重要,什么可以暂缓,哪些代价是我愿意承担的?
排序不会让选择变得轻松,也不能保证结果令人满意。它能做的是让决定更接近自己愿意认领的理由。过一段时间,结果可能改变我们对当初选择的看法;但至少,我们可以回头检查,当时是否认真考虑过自己的价值、现实条件和可能后果。
这不是要求每个人坚持同一套优先级。对一个人重要的事情,未必对另一个人同样重要。主体性也不是把偏好说得足够响亮,而是在不确定中辨认冲突,再决定什么暂时排在前面。
3. 责任需要边界
拥有判断,不等于要替所有人解决问题。责任、权限和资源经常分散在不同的人之间。若把所有责任都揽到自己身上,表面上像是在主动承担,实际上可能模糊了谁有权决定、谁能提供资源、谁应对结果负责。
我会用一个简单的问题检查这条边界:这件事是我的责任,还是我正在替别人承担?如果我没有相应的权限或资源,应该由谁参与?这不是把问题推开,而是让责任回到能够处理它的位置。
关系中的责任也不总能从一句承诺判断。持续的行动当然是重要信息,但单个行为不足以解释一个人的全部动机,也不足以替代沟通。我们可以观察对方做了什么,同时承认自己并不能仅凭观察读出对方的内心。
4. 关系不是自主的反面
自主不等于孤立。人与他人互相影响,也会接受支持、协商安排、依赖专业服务,或承担共同责任。这些连接本身并不取消选择能力。
更值得检查的是:价值排序由谁决定?决定由谁作出?后果由谁承担?当答案长期不一致时,关系可能让人失去行动空间;当各方能表达需求、协商责任并保留选择,连接也可能成为行动的条件。
依赖有时是理性的。向可信任的人求助、使用公共服务、与团队合作,不意味着把人生方向交出去。相反,知道自己的能力边界并选择合适的支持,可能正是负责任的判断。结构性限制也真实存在:并非所有人都拥有同样的时间、收入、安全感或退出选项。因此,不能把每一种困境都解释成个人不够独立。
5. 行动让判断落地
价值排序若始终停留在想法里,很难帮助我们辨认自己正在承担什么。行动可以小到一次对话、一项明确的请求、一段需要完成的工作,或一个暂时保留选择的决定。它不必证明一个人已经自信,也不必立刻改变全部处境。
行动会带来新的信息:我是否低估了成本?是否需要帮助?原先的优先级还适合现在吗?这些反馈可以促使我们调整,而不是把最初的决定当成必须维护的身份。
我会把主体性落到三个问题上:
- 此刻什么最重要?
- 哪一部分确实由我负责?
- 我愿意认领什么行动,也愿意承担它带来的哪些后果?
答案不一定一次就清楚。能够继续修正,也属于判断的一部分。
6. 主体性有边界
主体性不是“所有问题都靠自己解决”,也不是把每一个结果都归咎于个人选择。现实条件会限制可选路径;求助、合作和依赖可以是合理选择;责任也有边界。一个人可以努力行动,同时承认有些事情并不由自己控制。
因此,我不把自主理解成永不需要别人,而更看重能否说清自己的理由、辨认责任归属,并在可行范围内作出愿意承担的选择。
结语
主体性不是离开人群,也不是拒绝关系。它是在关系、限制和不确定之中,仍然尝试看清自己重视什么、哪些责任属于自己,以及下一步愿意做什么。
我们不必靠孤立证明自主。更实际的做法,是在连接中保留判断,在需要时接受支持,并对自己真正作出的选择负责。
Writing Note:这篇文章最初受到钱婧老师《钱婧老师的会客厅》Vol.169 关于“主体性”的讨论启发。本文关于价值排序、责任边界与行动的组织和论述,是我在此基础上的个人综合与延展,不使用节目逐字引文。
Narrowing a Research Question: Structure Is Not Evidence
Keeping research questions, evidence, and conclusions at the same scale.
- Observation
- Evidence
- Inference
- Bounded Claim
Structure ≠ EvidenceCausal conclusions require additional identification design.
1. Research questions tend to grow
A question may begin with a simple curiosity: a technology is changing work, people are adopting a new process, or an organization is trying a different governance arrangement. It is tempting to keep adding variables, mechanisms, and hypotheses until the model seems to include everything relevant.
The model can become more complete without the question becoming clearer. Every added concept needs a definition, every proposed relationship needs evidence, and every outcome needs an observable measure. If those requirements do not keep pace, structure can hide uncertainty rather than resolve it.
In my own research preparation, I have repeatedly needed to remind myself that documents, theories, and frameworks can become neatly organized before the evidence is sufficient. A question must move beyond “What seems worth explaining?” to “What can I observe, and how far do those observations let me go?”
2. Separate observation, inference, and causal claims
Three levels are particularly easy to collapse.
Observation
What does the source directly show? A public document, for example, may describe a policy or process. The careful first step is to state what the source disclosed, when it did so, and what its scope covers.
Mechanism or capability inference
What might those observations mean? A set of arrangements or actions over time may offer clues about a mechanism or capability. A clue remains an inference. The researcher must explain why the material supports that interpretation and whether other explanations remain plausible.
Causal claim
Claiming that an arrangement caused an outcome raises the evidential bar. The design needs to support time order, address alternatives, and establish what changed. A disclosure, a fragment from one case, or a theoretically plausible mechanism cannot by itself prove a causal effect.
These levels are not interchangeable. Public disclosure first tells us what an organization said publicly; it does not automatically establish a real and sustained organizational capability. The absence of disclosure does not prove that something is absent either. The strength of a conclusion has to stay within what the material can support.
3. What a polished model can conceal
A well-structured model can still contain several problems. A concept may be ambiguous. The evidence may not make a key variable observable. A proposed mechanism may lack supporting material. Co-occurring changes may be described as causal effects. Or the chosen measure may fail to capture the construct the researcher actually cares about.
These problems are related, but none substitutes for another. Drawing a variable in a diagram does not make it measurable. Finding an indicator does not establish that the measure is valid. Seeing two things occur together does not identify a causal relationship.
4. Narrowing does not diminish the research
A narrower question may appear to make a smaller contribution. It can also bring the question, materials, and conclusion onto the same scale. Instead of claiming that a broad mechanism has improved overall organizational decision quality, a study might first examine which divisions of work, review arrangements, authorization steps, or process changes are actually visible in public material—provided those details can be located and checked.
A narrower claim is not automatically valuable, and it does not guarantee a sound design. Its advantage is that readers can see what the sources record, where the researcher's interpretation begins, and which questions remain unanswered.
5. The questions I use now
Before adding another theory, variable, or hypothesis, I ask:
- What exactly am I claiming?
- What evidence would support that claim?
- What evidence do I have, or can I realistically collect?
- Which step is still my inference?
- What should I not claim yet?
If I cannot answer clearly, the next step is usually not another layer of structure. It is to narrow the question, check whether the evidence is available, or acknowledge what remains unresolved.
6. What this check does not solve
Narrowing a question does not automatically make a study good. A fit between question and evidence does not establish causal identification. A clear concept does not guarantee valid measurement. Researchers still need to assess source quality, measurement choices, analytical procedures, and alternative explanations.
The purpose of this check is more limited: it reminds me not to let conclusions run ahead of evidence. It is a constraint on research judgment, not a formula that guarantees quality.
Closing
Research design is not the task of fitting every plausible element into one model. It is the task of keeping the question, evidence, and conclusion on the same scale.
When a structure starts to look increasingly complete, I step back and ask: What do the sources actually show? Where does interpretation begin? Has the conclusion I am preparing to make crossed the boundary of what those sources can support?
1. 研究问题很容易越做越大
一个现象起初可能只是一个好奇:某种技术正在改变工作,人们开始采用新的流程,或者组织正在尝试一种治理安排。接下来,很容易不断加入变量、机制和假设,希望把所有看起来相关的因素都放进模型。
模型变得更完整,研究问题却未必更清楚。新增的每一个概念都需要定义,每一种关系都需要证据,每一个结果变量都需要可观察的指标。如果这些要求没有跟上,结构只会把不确定性藏得更深。
我在自己的研究准备中反复遇到一个提醒:文件、理论和框架可以很快变得齐全,但它们的齐全并不证明研究已经有足够证据。研究问题需要从“我觉得值得解释什么”继续收窄到“我能观察什么,以及观察结果允许我说到哪一步”。
2. 把观察、推断和因果主张分开
有三层内容尤其容易混在一起:
观察
来源直接显示了什么?例如,一份公开文件描述了某项制度或流程。此时最稳妥的表述是:该来源披露了什么、在什么时间披露、覆盖什么范围。
机制或能力推断
这些观察可能意味着什么?多项流程安排或跨时间的行动,或许构成某种机制或能力的线索。但“线索”仍是推断,需要解释为什么这些材料支持该解释,也要说明其他解释是否可能。
因果主张
如果要说某种安排导致了某种结果,证据要求会进一步提高。需要能支持时间顺序、替代解释和结果变化的研究设计。单一披露、一个案例片段或理论上合理的机制,不能独自证明因果效果。
这三层不是同义词。公开披露首先说明组织公开说了什么,不自动等于真实、持续的组织能力;没有披露某件事,也不等于它不存在。结论的强度必须受资料实际能支持的范围约束。
3. 漂亮模型可能藏着什么
结构完整的模型仍可能有几类问题。概念也许定义含混;证据可能无法观察模型中的关键变量;假定的机制可能没有材料支撑;相关变化可能被写成因果影响;而测量方式也可能没有捕捉到研究者真正关心的构念。
这些问题彼此相关,却不能互相替代。把变量画进图里,不会让它自动变得可测;找到一个观察指标,也不意味着测量有效;看到两个现象同时出现,也不等于识别了因果关系。
4. 收窄不是削弱研究
收窄问题,有时会让贡献看起来更小,但它能让问题、资料与结论使用同一尺度。与其声称一个宏大的机制已经提高了组织整体决策质量,不如先研究公开资料实际显示了哪些分工、复核、授权或流程调整——前提是这些安排确实可观察、可定位、可复核。
较窄的主张并不自动更有价值,也不保证研究设计正确。它的优势在于让读者看得清:哪些是来源记录,哪些是研究者的解释,还有哪些问题暂时没有被回答。
5. 我现在使用的检查方式
在继续增加理论、变量或假设之前,我会先问:
- 我到底在主张什么?
- 哪种证据能够支持这项主张?
- 我目前拥有什么证据,或实际能够收集什么?
- 哪一步仍然是我的推断?
- 现在有哪些内容还不应该主张?
如果回答不清楚,下一步通常不是再加一层结构,而是收窄问题、检查证据可得性,或承认当前仍有未解决的部分。
6. 这套检查没有解决什么
研究问题收窄,不等于研究自然就好。证据与问题匹配,不等于因果识别已经成立。概念定义清楚,也不等于测量有效。研究者仍然需要检查资料质量、测量选择、分析程序和替代解释。
这套检查的作用更有限:提醒我不要让结论跑到证据前面。它是一种研究判断的约束,不是保证研究质量的公式。
结语
研究设计不是把所有合理的元素都装进一个模型,而是让问题、证据与结论保持同一个尺度。
当结构看起来越来越完整时,我会退一步再问:来源实际显示了什么?解释从哪里开始?我准备说出的结论,是否已经越过了这两者能够支持的边界?
A small, curated knowledge layer.
Knowledge objects are concepts, frameworks, and research notes that I continue to connect and develop. This is an early public layer, not a complete digital garden.
Human-in-the-Loop
Insert an approval gate at decision points carrying high risk or irreversibility, trading low-friction human confirmation for controllable automation.
Organizational AI Governance
An organisational system of rules, practices, processes, and technical tools that keeps organisational use of AI aligned with strategy, goals, values, legal requirements, and ethical principles.
Retrieval Evaluation
A method for evaluating a retrieval system: paired comparison under a fixed corpus, fixed annotation set, and fixed evaluator, to isolate the increment from architecture change — with 'what counts as passing' written as layered hard gates rather than an aggregate score.
Evidence-Gated Evaluation
Separates development from independent admission evidence; holdout and admission remain incomplete.
Redistributing Agency in AI-Native Work
An author-developed lens on goals, constraints, verification, and commitment as AI takes on execution.
Base Rates and Judgment Calibration
Combines comparable-case information with case-specific evidence while making source limits visible.
The private Vault does not automatically sync to this website. Public knowledge is selectively reviewed, edited, and released.
Human-in-the-Loop
Insert an approval gate at decision points carrying high risk or irreversibility, trading low-friction human confirmation for controllable automation.
Why it matters: it makes the operational question visible — who may stop a process, and when. The concept describes a control point; it does not establish that any particular gate is effective.
Scope boundary: use this as a design concept for consequential decisions. It is not a claim that every action needs human approval or that a human check guarantees correctness.
Organizational AI Governance
An organisational system of rules, practices, processes, and technical tools that keeps organisational use of AI aligned with strategy, goals, values, legal requirements, and ethical principles.
Why it matters: it frames AI governance as coordinated organisational work across rules, practices, processes, and tools, rather than as a tool setting alone.
Evidence boundary: this is an author definition. It is not presented as an established theory of organisational capability, nor as evidence that a specific governance arrangement produces better outcomes.
A Retrieval Evaluation Framework Tested Within a Personal Knowledge Base
A paired evaluation method that holds the corpus, annotation set, and evaluator constant to isolate the increment from a retrieval-architecture change. Passing conditions are expressed as layered gates rather than a single aggregate score.
Why it matters: if material or evaluation conditions change between runs, a score difference cannot be attributed to architecture alone.
Scope boundary: tested within one personal knowledge base, with context-specific results. This is not a general benchmark; no private corpus, evaluation item, or score is shown here.
Evidence-Gated Evaluation: Independent Holdouts and Admission
- Development and iteration
- Freeze candidate
- Evidence boundary · no reuse of development evidence
- Independent holdout
- Hard-gate review
- Admission remains on HOLD
Protocol defined; holdout and admission are not completed.
The problem
An evaluation can become less convincing when the same evidence is repeatedly used to tune a system and to demonstrate that the system works. Once developers have seen the cases, their choices may adapt—deliberately or otherwise—to those cases. A strong result on familiar evidence can therefore describe progress on the development set without establishing how the candidate performs on independent demands.
This is an evaluation-design problem. It does not imply that every reused test is useless; it means the strength of a conclusion depends on what the evidence was allowed to influence.
Development evidence and admission evidence
Development evidence helps a team find errors, compare iterations, and improve a candidate. Admission evidence asks a different question: after the candidate and evaluation conditions have been fixed, does the system meet the conditions required for a higher level of use?
The distinction is about the role of evidence, not the name of a dataset. Evidence used to guide iteration should not silently become independent confirmation of the same iteration.
Why holdout matters
A holdout is evidence reserved from development decisions and evaluated only after the candidate and relevant evaluation conditions are frozen. Its purpose is to reduce contamination between improvement and confirmation.
A holdout does not guarantee generalisation. Its value depends on case independence, scope coverage, and a defensible evaluation process. A small or unrepresentative holdout can still support only a narrow conclusion.
Why hard gates matter
In this admission framework, some conditions are intentionally non-compensatory. A high aggregate score cannot compensate for a failed critical gate, a validity problem, or a boundary violation. When a condition is essential to admission, it should remain visible as a gate rather than disappear inside a composite score.
This framework therefore treats admission as a sequence of evidence checks, not a contest to maximise one blended number.
Layered evidence
A useful mental model is:
component evidence → integrated evidence → independent holdout evidence → admission decision
Each layer answers a different question. Evidence about a component does not by itself establish that the integrated system behaves acceptably. Evidence from an integrated run does not by itself establish performance on cases withheld from development. The admission decision must state which layers were actually completed.
What this framework does not prove
The protocol behind this page defines an evaluation design, holdout requirements, and hard-gate logic. It does not establish that a holdout dataset has been completed, that a blind evaluation has been run, or that an AI system has passed admission.
Current status: protocol defined; holdout not completed; admission remains on HOLD. No admission result or system-effectiveness claim is made here.
Relation to retrieval evaluation
A retrieval evaluation framework asks: How should retrieval changes be compared under controlled conditions?
Evidence-gated evaluation asks: When is the evidence strong enough to admit a candidate to a higher validation level?
The questions are related, but they are not interchangeable. A controlled comparison can inform development; it does not automatically replace independent admission evidence.
Claim and evidence notes
| Claim | Evidence basis | Strength and wording boundary |
|---|---|---|
| The protocol defines separate development and admission roles for evidence. | Designed protocol. | Method-design claim only. |
| The protocol defines a holdout requirement and layered hard gates. | Designed protocol. | Does not imply execution or successful outcomes. |
| Holdout can reduce evidence contamination. | Evaluation-design rationale. | Say “helps reduce”; do not claim it eliminates bias or proves generalisation. |
| The current system has passed admission. | No supporting evidence; holdout not completed. | Do not make this claim. Admission remains on HOLD. |
| The system is effective or improved. | No independent admission result established here. | Do not use “validated,” “passed,” “admitted,” or “proven improvement.” |
问题
当同一批证据反复用于调优系统、又用于证明系统有效时,评估结论会变得不够独立。开发者已经看过测试案例,后续选择可能有意或无意地适应这些案例。因此,熟悉材料上的好结果,可以说明系统在开发证据上有所进展,却不能单独说明它面对独立需求时表现如何。
这是评估设计问题,并不意味着所有重复使用测试都毫无价值;它意味着结论的强度取决于这些证据曾影响过什么决策。
开发证据与准入证据
开发证据帮助团队发现错误、比较迭代并改进候选系统。准入证据则回答另一个问题:在候选系统及评价条件固定之后,它是否满足进入更高使用或验证层级的条件?
两者的区别在于证据承担的角色,而不只是数据集名称。用于指导迭代的证据,不应不加说明地变成对同一轮迭代的独立确认。
为什么需要 Holdout
Holdout 是一组不参与开发决策、只有在候选系统和相关评估条件冻结后才用于评价的证据。它的作用是降低“改进系统”与“确认系统”之间的证据污染。
Holdout 不保证泛化。其价值仍取决于案例独立性、范围覆盖和评价过程是否站得住脚。规模很小或缺乏代表性的 holdout,仍只能支撑有限结论。
为什么需要硬门槛
在这套准入框架中,有些条件被有意设计为不可补偿。较高的综合分不能抵消关键门槛失败、评价效度问题或边界违规。如果某项条件是准入所必需的,它就应作为清晰可见的门槛,而不是被综合分掩盖。
因此,这个框架把准入视为一组分层证据检查,而不是最大化一个混合分数。
分层证据
可以用下面的心智模型理解:
组件证据 → 集成证据 → 独立 Holdout 证据 → 准入决策
每一层回答不同问题。组件表现不能单独证明集成系统的行为可接受;集成运行的证据也不能单独证明系统在开发阶段未见案例上的表现。准入结论必须说明哪些层级确实完成了。
这个框架不能证明什么
本页所依据的协议定义了评估设计、holdout 要求和硬门槛逻辑。它并未证明 holdout 数据集已经完成、盲评已经执行,或某个 AI 系统已经通过准入。
当前状态:协议已定义 → Holdout 尚未完成 → 准入仍为 HOLD。本页不声称已经取得准入结果,也不声称系统效果已获证明。
与检索评估的关系
检索评估框架回答:如何在受控条件下比较检索改动?
证据门控评估回答:证据何时足以支持候选系统进入更高验证层级?
两者相关,但不能互相替代。受控比较可以支持开发判断,却不会自动取代独立准入证据。
主张与证据说明
| 主张 | 证据基础 | 强度与措辞边界 |
|---|---|---|
| 协议区分开发证据和准入证据的角色。 | 已设计的协议。 | 仅为方法设计主张。 |
| 协议定义了 holdout 要求和分层硬门槛。 | 已设计的协议。 | 不代表已执行或已有成功结果。 |
| Holdout 有助于降低证据污染。 | 评估设计理由。 | 可说“有助于降低”,不可说消除偏差或证明泛化。 |
| 当前系统已经通过准入。 | 无支持证据;holdout 尚未完成。 | 不得提出此主张;准入仍为 HOLD。 |
| 系统有效或已经改进。 | 本对象没有建立独立准入结果。 | 不使用“已验证”“已通过”“已准入”或“改进已获证明”。 |
Redistributing Agency in AI-Native Work
Human
- Goal setting
- Constraint design
- Verification architecture
- Accountability and commitment
Harness
Rules · routing · checks · escalation
AI
Generation · search · transformation · delegated tasks
Execution may move toward AI; decision authority and accountability do not automatically move with it.
The shift
AI systems can increasingly execute tasks that previously required direct human effort. That change raises a question beyond task replacement:
As AI takes on more execution, which forms of agency should remain human?
Here, agency means the ability and authority to shape goals, set constraints, judge whether work is acceptable, and take responsibility for consequential commitments. This page offers an organising framework for thinking about those roles; it is not a validated general theory of organisational design.
Execution is not decision authority
Delegating execution does not automatically delegate the authority to decide what should be done, what risks are acceptable, or who is accountable when an outcome matters.
An AI system may draft, search, transform, or route information. A human or organisation still has to define the objective, establish boundaries, decide how outputs will be checked, and determine which commitments require human ownership. The appropriate allocation depends on the task, its uncertainty, and the consequences of error.
Four dimensions of retained human agency
The source synthesis distinguishes four human roles:
- Goal setting — defining objectives and priorities.
- Constraint design — specifying permissions, boundaries, risk limits, and acceptance criteria.
- Verification architecture — deciding what evidence and checks are needed, and when exceptions should be escalated.
- Accountability and commitment — owning consequential decisions, especially when a choice is difficult to reverse.
These are lenses for examining a workflow, not a universal sequence or a claim that every decision must remain human. Their boundaries can overlap, and actual allocations require context-specific design.
What should remain human?
A useful design question is not “Can this task be automated?” alone. Ask also:
- Who chooses the objective and resolves conflicts between objectives?
- Who sets the boundary of acceptable action?
- What evidence is enough to trust the output?
- Which errors require escalation?
- Who can pause, reverse, or own a consequential commitment?
As execution moves toward AI, accountability does not move with it by default. That is a governance choice that should be made explicitly.
Why harnesses matter
A harness is the surrounding structure that constrains, routes, and verifies AI execution. It can make a workflow more legible by stating its inputs, limits, checks, and escalation points.
A harness does not make uncertain goals clear by itself, guarantee that verification is reliable, or remove human responsibility. It is a way to organise delegation and oversight, not a substitute for judgment.
Open questions
The framework motivates questions that remain open rather than established conclusions:
- Under what conditions does redistributing execution change decision quality?
- When does a harness reduce coordination burden, and when does it add overhead?
- How does the right allocation change as task uncertainty or reversibility changes?
- Which forms of verification can be delegated without weakening accountability?
No productivity, scalability, cost-saving, or decision-quality effect is claimed here.
Relation to Human-in-the-Loop
Human-in-the-Loop asks when human intervention or approval is needed at a decision point.
Agency redistribution asks how authority and responsibility are allocated across a workflow, including before and after that intervention point.
The first focuses on an intervention mechanism; the second is a broader lens for mapping decision roles.
Claim and evidence notes
| Claim | Evidence basis | Strength and wording boundary |
|---|---|---|
| The four roles organise the source's account of agency. | Author's conceptual synthesis. | Present as this author's framework, not an established taxonomy. |
| A harness constrains, routes, and verifies AI execution. | Source definition and design synthesis. | Explanatory definition, not evidence of improved outcomes. |
| Execution can be delegated without automatically transferring accountability. | Conceptual distinction in the source. | A design principle/question, not a measured universal effect. |
| Harnesses improve productivity, scalability, or decision quality. | Not established by the source. | Keep as open questions; make no outcome claim. |
| AI can replace an entire company or team. | Not established and outside scope. | Do not make this claim. |
转变
AI 系统正在承担越来越多过去需要人直接完成的任务。这带来的问题不只是岗位替代:
当 AI 承担更多执行时,哪些形式的能动性应当留在人类手中?
这里的 Agency(能动性)指影响目标、设定约束、判断工作是否可接受,并对重要承诺承担责任的能力与权力。本页提供一个组织这些问题的分析框架;它不是已经验证的一般组织设计理论。
执行权不等于决策权
把执行任务交给 AI,并不会自动把“决定做什么”“哪些风险可以接受”或“重要结果由谁负责”的权力一并交出去。
AI 可以起草、检索、转换或路由信息;人或组织仍需定义目标、设置边界、决定如何检查输出,并判断哪些承诺需要由人负责。合理分工取决于任务、不确定性和错误后果。
人类 Agency 的四个维度
来源综合将人的角色组织为四个维度:
- 目标设定:定义目标与优先级。
- 约束设计:说明权限、边界、风险限制和验收标准。
- 验证架构:决定需要哪些证据与检查,以及何时升级异常。
- 责任与承诺:对重要决策负责,尤其是难以逆转的选择。
这些层面是分析工作流的视角,不是普遍适用的固定顺序,也不意味着每项决策都必须由人完成。它们可能彼此交叠,实际分配需要结合情境设计。
哪些权力应当留给人?
设计时不能只问“这项任务能否自动化”,还应问:
- 谁选择目标,并处理目标之间的冲突?
- 谁划定可接受行动的边界?
- 什么证据足以支持对输出的信任?
- 哪些错误必须升级处理?
- 谁能暂停、撤回或承担重要承诺?
执行权向 AI 转移时,责任不会自动随之转移。这是一项需要明确作出的治理选择。
Harness 为什么重要
Harness 是围绕 AI 执行建立的结构,用来约束、路由和验证任务。它可以通过明确输入、限制、检查与升级节点,使工作流更清晰。
Harness 本身不会自动澄清不确定目标,不保证验证可靠,也不会消除人的责任。它用于组织委派与监督,不能替代判断。
尚待回答的问题
这个框架引出一些仍待研究的问题,而不是已经成立的结论:
- 在什么条件下,重新分配执行权会改变决策质量?
- Harness 何时降低协调负担,何时反而增加开销?
- 任务不确定性或可逆性变化时,合适的分工如何变化?
- 哪些验证工作可以委派而不削弱责任边界?
本页不声称生产率、组织扩展性、成本节约或决策质量已经改善。
与 Human-in-the-Loop 的关系
Human-in-the-Loop 回答:在某个决策点,何时需要人工介入或批准?
Agency 再分配回答:一项工作流中,决策权和责任如何分布?这也包括介入点之前与之后的分工。
前者关注介入机制;后者提供一张更广的决策角色地图。
主张与证据说明
| 主张 | 证据基础 | 强度与措辞边界 |
|---|---|---|
| 四个角色概括了来源中对 Agency 的组织方式。 | 作者的概念综合。 | 作为作者框架呈现,不称为已确立的分类法。 |
| Harness 用于约束、路由和验证 AI 执行。 | 来源定义与设计综合。 | 属于解释性定义,不是效果证据。 |
| 执行可以委派,但责任不会自动转移。 | 来源中的概念区分。 | 作为设计原则与问题提出,不称为普遍测量结果。 |
| Harness 提高生产率、扩展性或决策质量。 | 来源未建立这些结果。 | 仅列为开放问题,不作效果主张。 |
| AI 可以替代整个公司或团队。 | 来源未建立,且超出本页范围。 | 不提出此主张。 |
Base Rates and Judgment Calibration
Review alone does not establish causality or method effectiveness.
One vivid case, or a set of comparable cases?
A plan can look carefully reasoned. The team explains its assumptions, lays out the steps, and describes why this situation is different. Before accepting that account, it helps to ask a quieter question: how often have comparable plans reached the intended outcome?
That question moves judgment from the case in front of us to the broader set of cases it belongs to. A base rate provides context from that set; case-specific information helps us decide whether this case differs in a relevant way. They play different roles, and neither automatically replaces the other.
Start with the reference class, then return to the case
Kahneman and Tversky’s research on prediction discussed the representativeness heuristic: people may form predictions from how much a case resembles a familiar prototype, without adequately considering prior probabilities or the reliability of evidence. Their later review of judgment under uncertainty described representativeness, availability, and anchoring as heuristics that are often useful but can also produce systematic biases.[1][2]
This does not mean that base rates are always right or that case details do not matter. It suggests a check: when a story feels especially vivid, coherent, or persuasive, ask what reference group makes it representative. Then examine whether the supporting evidence is reliable and genuinely diagnostic.
Put both kinds of information on the table
Prediction can involve information about an individual case and information about a distribution of cases. In their discussion of intuitive prediction and corrective procedures, Kahneman and Tversky distinguish data about a singular case from distributional data.[3]
In the narrower domain of project forecasting, reference-class forecasting offers an outside view: examine the actual outcomes of comparable projects, then return to the project at hand. Flyvbjerg’s account focuses on project forecasting, particularly cost risk. The point used here is the practice of consulting relevant past cases; the method’s effects are not generalized to every personal or organizational judgment.[4]
In practice, ask:
- What kind of judgment is being made?
- Which past cases are sufficiently comparable on the characteristics and outcome conditions relevant to this judgment?
- What context does their outcome distribution provide?
- What information about the current case is credible, relevant, and diagnostic?
- Are the differences from the reference cases strong enough to justify a different judgment?
This set of questions is the author’s working synthesis, not a verbatim procedure from a single study.
Choosing a reference class is itself a judgment
Selecting a comparison group is not mechanical. A class that is too broad may combine unlike cases; one that is too narrow may leave too little data or include only cases that fit an expectation. Changes in organizational, technological, or institutional conditions can also limit how well an older distribution applies.
A base rate is therefore an input whose source, scope, and limitations need to be explained, not an answer that makes the decision for us. Case-specific evidence deserves the same scrutiny: Is its source reliable? Is it relevant? Is it strong enough to change the judgment?
Record the judgment and leave room to review it
One practical discipline is to record, before the outcome is known, the reference class, the base-rate information used, the key case-specific evidence, the forecast, and its uncertainties. Review the judgment against the outcome later. Keeping a record of forecasts and outcomes is a practical recommendation from the author; it should not be mistaken for a universally validated intervention established by the sources above.
Calibration here primarily means making a forecast traceable and comparing it with the eventual outcome. Unless calibration metrics are actually computed, the term does not imply formal probabilistic calibration.
AI can help organize comparable cases, make assumptions visible, or generate alternative explanations. A fluent case narrative still needs checks on its sources, relevance, and scope. This is a conceptual connection to the judgment process; it does not claim that AI causes base-rate neglect or that AI use necessarily improves forecasting.
Limits
- The cases found may not form an appropriate reference class; comparability needs an argument.
- Small samples, data quality, and selection methods affect how a distribution should be interpreted.
- Structural change or a genuinely rare new situation may reduce the relevance of historical comparisons.
- This essay concerns prediction and judgment calibration. It does not infer causality or provide a statistical model or individualized probability advice.
- How base rates and case-specific evidence should be weighed depends on the question, evidence quality, and context.
Claim and evidence boundaries
| Statement | Evidence type | Boundary |
|---|---|---|
| Representativeness-based judgment may underweight prior probability or evidence reliability | Supported by research sources: [1] | Describes tendencies in particular judgment research, not every person or every judgment |
| People use several heuristics under uncertainty, which can produce systematic biases | Supported by research source: [2] | Does not mean heuristics are always wrong |
| Case-specific and distributional information are distinct inputs to prediction | Original authors’ methodological discussion: [3] | Not presented as a fixed formula for every situation |
| Reference-class forecasting can use actual outcomes of comparable projects to inform project forecasts | Domain-specific method source: [4] | Limited to project forecasting; not generalized as a universal causal effect |
| List the reference class, base-rate context, case evidence, and review the forecast after the outcome | Author synthesis and practical recommendation | Not a complete process directly tested by one cited source |
| Check the sources, relevance, and limits of AI-generated explanations | Author’s process reminder | Does not claim AI causes base-rate neglect or improves calibration |
References
- Kahneman, D., & Tversky, A. (1973). “On the Psychology of Prediction.” Psychological Review, 80(4), 237–251. https://doi.org/10.1037/h0034747
- Tversky, A., & Kahneman, D. (1974). “Judgment under Uncertainty: Heuristics and Biases.” Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124. PubMed record
- Kahneman, D., & Tversky, A. (1982). “Intuitive Prediction: Biases and Corrective Procedures.” In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment under Uncertainty: Heuristics and Biases, 414–421. Cambridge University Press. https://doi.org/10.1017/CBO9780511809477.031
- Flyvbjerg, B. (2006). “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, 37(3), 5–15. https://doi.org/10.1177/875697280603700302
一个具体案例,还是一组相似案例?
一个计划看起来很周全:团队解释了关键假设,列出执行步骤,也说明了为什么这次情况不同。这样的叙述可能很有说服力。但在接受它之前,还可以问一个更朴素的问题:过去相似的计划通常走到了哪一步?
这个问题把判断从眼前个案带到它所属的一组案例。基准率提供这组案例的背景;个案信息则帮助我们判断这一次是否确有重要差异。二者承担不同角色,不能由其中一方自动取代另一方。
先看参照类,再回到个案
Kahneman 与 Tversky 对预测判断的研究讨论了代表性启发式:人们可能依据一个案例与某种典型情形有多相似来形成预测,而没有充分考虑先验概率或证据可靠性。后续关于不确定情境下判断的综述也指出,代表性、可得性和锚定等启发式通常有用,却可能带来系统性偏差。[1][2]
这并不意味着基准率永远正确,也不意味着个案细节都不重要。它提示的是一种检查方式:当一个故事格外鲜明、连贯或令人信服时,先问它相对于什么参照群体显得典型;再检查支持它的证据是否可靠、是否真正具有区分力。
把两类信息放在同一张桌上
预测可以同时面对个案层面的信息与分布层面的信息。Kahneman 与 Tversky 对直觉预测与校正程序的讨论明确区分了关于单一案例的数据和关于一组案例分布的数据。[3]
在项目预测这一较窄的领域,reference-class forecasting(参考类预测)提供了一个外部视角:先观察可比项目的实际结果,再回到当前项目评估预测。Flyvbjerg 对这一方法的论述聚焦于项目预测,尤其是项目成本风险;这里借用的是它“先看相关案例的实际表现”的思路,不把该方法的效果外推到所有个人或组织判断。[4]
实际使用时,可以依次问:
- 这次判断属于哪一类问题?
- 哪些既往案例在与当前判断相关的关键特征和结果条件上足够可比,可以作为参照?
- 这些案例的结果分布提供了什么背景?
- 当前个案有哪些可信、相关且有区分力的信息?
- 当前情形与参照案例的差异,是否足以支持不同的判断?
这组问题是本文作者的工作性综合,不是对某一项研究程序的逐字转述。
参考类本身也需要判断
选择参照群体不是机械步骤。范围太宽,案例之间可能并不相似;范围太窄,样本可能不足,或者只挑进了与预期相符的案例。组织、技术或制度条件变化时,旧分布的适用性也需要重新评估。
所以,基准率不是替我们作决定的答案,而是一项需要说明来源、范围和局限的输入。个案证据也需要接受同样的检查:来源是否可靠,证据是否相关,它是否足以改变原有判断?
记录判断,留出复盘空间
一种实用做法是,在结果出现前记录:所选参照类、判断所依据的基准信息、当前个案的关键证据、预测及其不确定之处。之后再对照结果复盘。持续记录预测与结果,是本文提出的校准实践建议;它不应被误读为上述文献已验证的通用干预效果。
这里的“校准”主要指把事前判断与事后结果进行可追溯比较;除非实际计算校准指标,否则不表示统计意义上的概率校准。
AI 可以帮助整理可比案例、显化假设或生成替代解释。但一个流畅的个案故事仍需经过来源、相关性与适用范围的检查。这里讨论的是判断流程上的关联,并不主张 AI 会导致基准率忽视,也不声称使用 AI 必然改善预测。
适用边界
- 找到的案例未必构成合适的参考类;可比性需要论证。
- 小样本、数据质量和选择方式会影响分布信息的解释。
- 结构变化或真正罕见的新情形可能削弱历史参照的适用性。
- 本文讨论预测与判断校准,不据此推断因果关系,也不提供统计模型或个体化概率建议。
- 基准率与个案证据如何权衡,仍取决于问题、证据质量与情境。
Claim 与证据边界
| 本文表述 | 证据类型 | 边界 |
|---|---|---|
| 代表性判断可能未充分考虑先验概率或证据可靠性 | 研究来源支持:[1] | 描述特定判断研究中的倾向,不等于所有人的每次判断 |
| 不确定判断会使用若干启发式,且可能产生系统性偏差 | 研究来源支持:[2] | 不表示启发式总是错误 |
| 个案信息与分布信息可作为不同类型的预测输入 | 原作者方法讨论:[3] | 不把它扩写为任何情境下的固定公式 |
| 参考类预测可用可比项目的实际表现校准项目预测 | 领域方法来源:[4] | 限定在项目预测语境,不外推为普遍因果效果 |
| 先列参考类、基准信息、个案证据并在结果后复盘 | 作者综合与实践建议 | 不是单篇来源直接验证的完整流程 |
| AI 生成的解释也应核对来源、相关性与边界 | 作者提出的流程提醒 | 不声称 AI 导致基准率忽视或提升校准 |
参考资料
- Kahneman, D., & Tversky, A. (1973). “On the Psychology of Prediction.” Psychological Review, 80(4), 237–251. https://doi.org/10.1037/h0034747
- Tversky, A., & Kahneman, D. (1974). “Judgment under Uncertainty: Heuristics and Biases.” Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124. PubMed record
- Kahneman, D., & Tversky, A. (1982). “Intuitive Prediction: Biases and Corrective Procedures.” In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment under Uncertainty: Heuristics and Biases, 414–421. Cambridge University Press. https://doi.org/10.1017/CBO9780511809477.031
- Flyvbjerg, B. (2006). “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal, 37(3), 5–15. https://doi.org/10.1177/875697280603700302
Open to thoughtful conversations.
Open to conversations around AI, decision systems, organizational adaptation, and research collaboration.