LAB 10 · 一次行动怎样改变判断
先阅读保存结果和解释,再按本册步骤选择是否运行。在 Deepnote 阅读与运行 · 下载 Notebook
LAB10|一次行动怎样改变对世界的判断
这册实验接住 N49–N50 的记忆与开门情境。你会先看两个候选怎样解释相同历史,再选一项动作,让独立环境返回结果,更新判断并实际执行下一项开门动作。最后把历史拿走,观察当前画面相同时会丢失什么。
运行身份:确定性规则环境与有限候选推断。没有神经网络训练,没有机器人,没有真实门传感器。曲线均标明对应的实际运行;展示图不参与环境计算。CPU 即可运行,无网络、密钥或权重下载。
原基线云端复核(2026-09-30,保留)。 本册已在 Deepnote 独立内核完整执行:六个情节分支均收到独立环境反馈,两种记忆策略都实际采取动作并更新结果。这仍是规则环境实验,训练参数数为0。
保留基线:按钮环境+approach,历史相容不等于机制已分清

读左图:两行是按钮机制与接近机制,四列是动作,OPEN/CLOSED 是执行前的候选预测。同时走近并按按钮时,两行都为 OPEN;远处等待时都为 CLOSED。两项旧记录都成立,仍然无法区分它们。
读右图:先验各 0.5,前两次独立环境反馈不改变权重。第三次只走近、没有按按钮,按钮环境返回 CLOSED;接近候选被排除。第四次控制器按更新后的判断选择远处按按钮,实际环境返回 OPEN。金色区域对应一次有区分力的干预。纵轴只表示列出的两个候选内部权重,不表示现实全部可能性已经被穷尽。
| 步骤 | 实际动作 | 按钮候选预测 | 接近候选预测 | 本次环境返回 | 更新后按钮/接近权重 |
|---|---|---|---|---|---|
| 旧历史1 | 走近并按键 | 开 | 开 | 开 | 0.5 / 0.5 |
| 旧历史2 | 回到远处等待 | 关 | 关 | 关 | 0.5 / 0.5 |
| 新干预 | 只走近,不按键 | 关 | 开 | 关 | 1 / 0 |
| 下一步选择 | 远处按键 | 开 | 关 | 开 | 1 / 0 |
表中“1”成立在明确候选与无噪声假设内。对现实门,一次不开可能来自故障、延迟或未进入范围;不能拿本例替代条件核对。
环境的规则与任务
我们把门的位置观察简化为远/近,把门状态简化为开/关。按钮是可从远处执行的按钮脉冲,模拟遥控按钮;每一步足够等待门响应。门每一步重新判断触发条件,没有惯性、延迟、故障或持续打开的记忆。一次运行中隐藏机制固定:按钮环境仅在本步按键时开门;接近环境仅在本步处于近处时开门。
任务是“下一步观察到门开”,不是走过门。指定动作成本用于同样能开门时的选择:远处按键为 1,只走近为 1.2,两者都做为 2.2。这些是课程给定的相对成本,不是物理测量。策略先最大化候选内部的开门概率,再选成本最低者。
旧历史只含“两者都做”和“远处等待”。环境机制在初始化时由实验者选择,但控制器只收到 step(action) 回调;它看不到隐藏机制。你能阅读环境源码,是为了检查教学实验;不能把那里写的答案作为控制器输入。
三件事分别发生:预测、反馈与更新
predict(hypothesis, action) 是候选预测表,只收候选名称与动作。DoorEnvironment.step(action) 是独立规则环境,直接执行动作并返回观察,不调用预测器。update(belief, action, observation) 比较预测与环境返回,再调整权重。这样候选可以是错的,反馈也真的有机会改变控制器。
确定性情况下,更新就是把与观察冲突的候选权重置零,再将剩余权重归一化:新权重与“旧权重×候选是否预测对”成正比。两个候选对所有旧记录都预测对,所以旧权重不变;只有让预测分开的动作,才可能改变相对判断。
若两个候选都预测关门,环境却返回开门,代码报告 No candidate explains feedback。这是有用的失败:候选集合或条件假设出了问题。程序不会把全零权重强行换成某个胜者,也不会让模型自己生成一段视频来填补反馈。
只改一个变量:重复成功,还是取得新信息
默认 INTERVENTION = "approach"。保留环境、历史、先验和成本不变,仅把它改为 "both",重算同一册。
以下是 2026-10-01 UTC 的两次独立脚本执行,固定按钮机制、both → wait_far 历史、各0.5先验与成本,仅改变干预。网页先接住这两次已保存结果;你重跑后应读练习代码紧接输出的新运行编号。
保存运行:approach
运行编号:button__approach__20261001T025630.481808Z__396147cc
| 阶段 | 实际动作 | 按钮候选预测 | 接近候选预测 | 环境实际返回 | 更新后权重(按钮 / 接近) |
|---|---|---|---|---|---|
| 共同历史 | both |
开 | 开 | 近 / 开 | 0.5 / 0.5 |
| 共同历史 | wait_far |
关 | 关 | 远 / 关 | 0.5 / 0.5 |
| 本次干预 | approach |
关 | 开 | 近 / 关 | 1 / 0 |
| 下一动作(已执行) | press_far |
开 | 关 | 远 / 开 | 1 / 0 |
本次干预返回关门,候选熵减少 1 bit;下一动作是 press_far,独立环境实际返回开门。完成开门目标与取得区分信息分别检查。

保存运行:both
运行编号:button__both__20261001T025631.027377Z__372b5e95
| 阶段 | 实际动作 | 按钮候选预测 | 接近候选预测 | 环境实际返回 | 更新后权重(按钮 / 接近) |
|---|---|---|---|---|---|
| 共同历史 | both |
开 | 开 | 近 / 开 | 0.5 / 0.5 |
| 共同历史 | wait_far |
关 | 关 | 远 / 关 | 0.5 / 0.5 |
| 本次干预 | both |
开 | 开 | 近 / 开 | 0.5 / 0.5 |
| 下一动作(已执行) | both |
开 | 开 | 近 / 开 | 0.5 / 0.5 |
本次干预返回开门,候选熵减少 0 bit;下一动作是 both,独立环境实际返回开门。完成开门目标与取得区分信息分别检查。

独立阅读:等权先验的候选信息量

这里的信息量来自列出的候选,且只按本例无噪声规则计算。先验熵是 1 bit,能完全分开两候选的动作把它降为 0;共同预测的动作不改变它。信息量没有包括成本,也没有把当前开门目标自动变成探索目标。
再做一个有意义的独立对照:将干预改成 "press_far"。在按钮环境返回开,在接近环境返回关;这个动作也能区分候选。它说明区分力并不专属于“走近”这个动作,而取决于候选在该条件下是否作出不同预测。
当前观察一样,为什么下一动作可以不同

先读两幅历史曲线的最后一列:两条记录都回到“远处、关门”。再向前看“只走近”:按钮环境返回关门,接近环境返回开门。这段历史已区分列出的机制,因此保留历史权重时,前者只按按钮,后者只走近。丢弃历史时,控制器把权重重置为各 0.5,两种相同当前观察下都选择保守的“两者都做”。
这次对照的四条分支都在独立规则环境中实际执行。每种固定机制下先重放同一段动作与反馈历史,再仅改变下一步决策是否保留历史权重;没有把候选预测填成无记忆策略的验证结果。
| 固定环境机制 | 下一步决策保留什么 | 实际执行动作 | 给定动作成本 | 独立环境返回 |
|---|---|---|---|---|
| 按钮触发 | 之前的反馈历史 | 远处按键 | 1 | 开门 |
| 按钮触发 | 仅当前“远处、关门” | 走近并按键 | 2.2 | 开门 |
| 接近触发 | 之前的反馈历史 | 只走近 | 1.2 | 开门 |
| 接近触发 | 仅当前“远处、关门” | 走近并按键 | 2.2 | 开门 |
没有记忆也能成功,因为本例允许一个更费动作的共同成功方案。历史的收益体现为下一步减少不必要动作,不是预设“有记忆成功、没记忆必败”。这里只比较取得同样历史后的下一步成本,没有把取得历史的成本或记忆计算成本计入;它不证明先探索总比立刻开门合算。
这里保留的是“哪种开门机制更符合反馈”的证据。门本身每步重新响应,没有依赖过去动作的隐含开门状态;杯中是否有水、被遮挡物体在哪里,属于本章另一类需要历史来估计当前物理状态的问题。两种用途都说明眼前观察可能不够,但不能混作同一种机制。
运行与保存
Notebook 含环境与控制器全部代码,可在新内核顺序运行。脚本版将环境保存在独立的 door_environment.py,与 lab10_action_feedback.py 放在同一目录,然后运行:
python lab10_action_feedback.py --output results
这条命令保存独立的默认基线 trajectories.json 与三张 PNG。单次干预另用:
python lab10_action_feedback.py --truth button --intervention approach --output my_results
# 下一轮只改 --intervention both
每次单次干预创建含机制、动作、UTC 时间与唯一编号的子目录,保存 record.json、candidate-weights.png、result.md、report.html 和校验清单,不覆盖另一轮。打开该目录的 report.html 即可先读表图,再展开 JSON。JSON 每一步保留执行前预测、动作、真实返回、更新前后权重及下一步选择。基线实际运行了两种隐藏机制与三种干预的六条分支,并在两种机制下分别实际执行有历史与无历史策略,共四条记忆对照分支;默认课堂图使用按钮环境与只走近干预。代码另检查外部反馈确实改变选择,以及候选全失配时是否报告问题。
从本册回到 第八章:生成怎样进入行动;如果只是没听懂 N50,先读 为什么历史正确仍需要行动。本册不增加计分要求。完成后留下一个动作对照和三句话:原来哪些候选都解释得通,本次反馈排除了什么,下一步选择因此怎样改变。
自检与答案
问:模型自己预测了一个开门视频,能否当作观察更新候选?答:这仍是候选后果,不能充当独立环境返回。此处只有环境 step 的观察进入更新。
问:干预成功开门,但权重仍为各 0.5,是不是程序没有学习?答:如果动作让两种候选共同预测开门,反馈本来就没有区分力。当前任务成功与获得新机制信息不是同一个结论。
问:这里只有两种写好的规则,哪里变了?答:候选参数没有训练,环境也没有学习;控制器根据反馈改变了候选权重与下一动作。若要研究学得世界模型,下一项工作才是从经历拟合预测参数,并在独立条件上评估,而不能改名就算完成。
实现一:独立环境
这段环境代码不调用候选预测器。每一步只按固定隐藏规则返回观察。
查看可执行代码
"""LAB10 的独立规则环境;不调用候选模型、后验或动作选择器。
每一步足够等待门响应。按钮是一次脉冲,门每一步重新响应;本模型不包含
惯性、延迟或故障。位置是远/近两个离散值。隐藏机制在一次运行中固定。
"""
from dataclasses import dataclass
@dataclass(frozen=True)
class Observation:
position: str
door_open: bool
class DoorEnvironment:
def __init__(self, mechanism="button"):
if mechanism not in ("button", "proximity"):
raise ValueError("mechanism must be button or proximity")
self.__mechanism = mechanism
self.__position = "far"
self.__door_open = False
self.__time = 0
def observe(self):
return Observation(self.__position, self.__door_open)
def step(self, action):
# 独立模拟动作执行,不读取下面课程预测器的表。
if action == "press_far":
self.__position, pressed = "far", True
elif action == "approach":
self.__position, pressed = "near", False
elif action == "both":
self.__position, pressed = "near", True
elif action == "wait_far":
self.__position, pressed = "far", False
else:
raise ValueError("Unknown action: " + str(action))
if self.__mechanism == "button":
self.__door_open = bool(pressed)
else:
self.__door_open = self.__position == "near"
self.__time += 1
return self.observe()
实现二:候选、更新与动作选择
观察只通过 step 回调进入控制器;没有隐藏机制输入。优先读 predict、collect 与 controller_episode 三处。
查看可执行代码
"""LAB10:候选预测 → 干预 → 独立反馈 → 更新 → 下一动作。
运行: python lab10_action_feedback.py --output results
需要 Python 3.10+、numpy、matplotlib;同目录需 door_environment.py。
无网络、无权重下载、无模型训练。预测器和规则环境分别实现。
"""
import argparse
from dataclasses import asdict
from pathlib import Path
import json
import math
import numpy as np
import matplotlib
import matplotlib.pyplot as plt
HYPOTHESES = ("button", "proximity")
ACTIONS = ("both", "press_far", "approach", "wait_far")
COST = {"both": 2.2, "press_far": 1.0, "approach": 1.2, "wait_far": 0.0}
def predict(hypothesis, action):
"""候选模型的预测。只用显式候选和动作;不知道环境的隐藏机制。"""
button_predictions = {"both": True, "press_far": True,
"approach": False, "wait_far": False}
proximity_predictions = {"both": True, "press_far": False,
"approach": True, "wait_far": False}
return {"button": button_predictions,
"proximity": proximity_predictions}[hypothesis][action]
def update(belief, action, observation):
"""确定性候选下的 Bayes 更新;反馈矛盾时保留失败,不强行归一化。"""
weights = {h: p * float(predict(h, action) == observation.door_open)
for h, p in belief.items()}
total = sum(weights.values())
if total == 0:
raise ValueError("No candidate explains feedback; expand/check the model class.")
return {h: p / total for h, p in weights.items()}
def entropy(belief):
return -sum(p * math.log2(p) for p in belief.values() if p > 0)
def expected_information(belief, action):
"""在列出的候选/概率内计算动作能区分多少,不访问环境。"""
prior = entropy(belief)
expected_posterior = 0.0
for result in (False, True):
weights = {h: p for h, p in belief.items() if predict(h, action) == result}
probability = sum(weights.values())
if probability:
expected_posterior += probability * entropy(
{h: p / probability for h, p in weights.items()})
return prior - expected_posterior
def next_open_action(belief):
"""任务仅为下一步看到门开;先最大化成功概率,再最小化预设成本。
不是通过门、不是机器人控制;成本是教学指定的动作代价。
"""
def score(action):
success = sum(p * float(predict(h, action)) for h, p in belief.items())
return (success, -COST[action])
return max(ACTIONS, key=score)
def collect(step, belief, action, phase):
# 所有预测在动作执行之前记录;观察只来自独立 step 回调。
predicted = {h: predict(h, action) for h in HYPOTHESES}
before = dict(belief)
observation = step(action)
after = update(before, action, observation)
return after, {
"phase": phase, "action": action, "predictions": predicted,
"observation": asdict(observation), "belief_before": before,
"belief_after": after, "entropy_before_bits": entropy(before),
"entropy_after_bits": entropy(after),
}
def controller_episode(step, intervention="approach"):
"""控制器只有 step 回调,绝不接收 truth/mechanism 参数。"""
if intervention not in ("approach", "press_far", "both"):
raise ValueError("Choose approach, press_far or both")
belief, records = {"button": .5, "proximity": .5}, []
for action in ("both", "wait_far"):
belief, row = collect(step, belief, action, "shared_history")
records.append(row)
information = {a: expected_information(belief, a) for a in ACTIONS}
belief, row = collect(step, belief, intervention, "intervention")
records.append(row)
action = next_open_action(belief)
belief, row = collect(step, belief, action, "next_goal_action")
records.append(row)
return {"records": records, "candidate_information_bits": information,
"final_belief": belief, "next_action": action,
"goal_open_succeeded": records[-1]["observation"]["door_open"]}
def memory_controller(step, keep_history=True):
belief = {"button": .5, "proximity": .5}
records = []
for action in ("both", "wait_far", "approach", "wait_far"):
belief, row = collect(step, belief, action, "memory_history")
records.append(row)
current = records[-1]["observation"]
with_memory = next_open_action(belief)
without_memory = next_open_action({"button": .5, "proximity": .5})
# 两种策略各在独立环境中执行相同前缀,随后仅改变是否保留历史权重。
history_belief = dict(belief)
decision_belief = history_belief if keep_history else {"button": .5, "proximity": .5}
executed_action = next_open_action(decision_belief)
phase = "memory_goal_action" if keep_history else "observation_only_goal_action"
belief, feedback = collect(step, decision_belief, executed_action, phase)
return {"history": records, "same_current_observation": current,
"history_conditioned_belief": history_belief,
"decision_belief": decision_belief,
"executed_policy": "history_conditioned" if keep_history else "observation_only",
"executed_action": executed_action,
"with_memory_action": with_memory,
"without_memory_action": without_memory,
"with_memory_cost": COST[with_memory],
"without_memory_cost": COST[without_memory],
"actual_feedback": feedback,
"comparison_scope": "Different hidden mechanisms; identical current observation. "
"Both is a robust costly action when history is discarded."}
def verify_separation():
"""真实风险检查:控制器必须用外部返回值更新,而非拿自身预测当真相。"""
# 与默认环境不同的反馈:approach 回开门,必须保留 proximity。
scripted = iter([Observation("near", True), Observation("far", False),
Observation("near", True), Observation("near", True)])
calls = []
def external_step(action):
calls.append(action)
return next(scripted)
result = controller_episode(external_step, "approach")
assert result["final_belief"] == {"button": 0., "proximity": 1.}
assert calls == ["both", "wait_far", "approach", "approach"]
# 一个“什么都没做却开门”的反馈不属于任一候选,必须暴露问题。
try:
update({"button": .5, "proximity": .5}, "wait_far", Observation("far", True))
except ValueError:
pass
else:
raise AssertionError("Contradictory feedback was silently accepted")
return {"external_feedback_controls_update": True,
"candidate_failure_is_reported": True}
实现三:保存全部分支与图
以下代码从运行记录生成图,不使用预写终点或手动数组替代环境反馈。
查看可执行代码
def plot_weights(episode, ax, title):
"""Every stage/action/outcome/weight comes from this episode's record."""
records = episode["records"]
x = np.arange(len(records) + 1)
colors = {"button": "#003262", "proximity": "#c38b16"}
for hypothesis in HYPOTHESES:
values = [records[0]["belief_before"][hypothesis]] + [r["belief_after"][hypothesis] for r in records]
ax.plot(x, values, marker="o", lw=2.5, color=colors[hypothesis], label=hypothesis + " candidate")
phases = {"shared_history": "History", "intervention": "Intervention", "next_goal_action": "Next goal action"}
labels = ["Prior"] + [phases.get(r["phase"], r["phase"]) + "\n" + r["action"] + "\n" + ("OPEN" if r["observation"]["door_open"] else "CLOSED") for r in records]
ax.set_xticks(x, labels, fontsize=9)
for i, record in enumerate(records, 1):
if record["phase"] == "intervention":
ax.axvspan(i - .35, i + .35, alpha=.12, color="#c38b16")
ax.set_ylim(-.08, 1.15); ax.set_yticks([0, .5, 1])
ax.set_ylabel("Weight within the TWO listed candidates")
ax.set_title(title, fontweight="bold", pad=16)
ax.legend(loc="upper left", fontsize=9)
def intervention_markdown(episode):
"""Readable core result; predictions and external observations have separate columns."""
phases = {"shared_history": "共同历史", "intervention": "本次干预", "next_goal_action": "下一动作(已执行)"}
state = lambda value: "开" if value else "关"
lines = ["| 阶段 | 实际动作 | 按钮候选预测 | 接近候选预测 | 环境实际返回 | 更新后权重(按钮 / 接近) |", "|---|---|---|---|---|---|"]
for r in episode["records"]:
observation = ("近" if r["observation"]["position"] == "near" else "远") + " / " + state(r["observation"]["door_open"])
weights = " / ".join(f'{r["belief_after"][h]:g}' for h in HYPOTHESES)
lines.append(f'| {phases.get(r["phase"], r["phase"])} | `{r["action"]}` | {state(r["predictions"]["button"])} | {state(r["predictions"]["proximity"])} | {observation} | {weights} |')
intervention = next(r for r in episode["records"] if r["phase"] == "intervention")
gained = intervention["entropy_before_bits"] - intervention["entropy_after_bits"]
feedback = episode["records"][-1]["observation"]["door_open"]
lines.append(f'\n本次干预返回**{state(intervention["observation"]["door_open"])}门**,候选熵减少 **{gained:g} bit**;下一动作是 **`{episode["next_action"]}`**,独立环境实际返回**{state(feedback)}门**。完成开门目标与取得区分信息分别检查。')
return "\n".join(lines)
def present_intervention(episode, truth, intervention, output="my_results", display_result=True):
"""Save and display THIS run, without replacing baseline or another intervention."""
from datetime import datetime, timezone
from uuid import uuid4
import hashlib
import html
timestamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%S.%fZ")
run_id = truth + "__" + intervention + "__" + timestamp + "__" + uuid4().hex[:8]
folder = Path(output) / run_id
folder.mkdir(parents=True, exist_ok=False)
identity = {"run_id": run_id, "generated_at_utc": timestamp,
"kind": "actual_independent_rule_environment_execution", "truth_selected_by_experimenter": truth,
"intervention": intervention, "shared_history": [r["action"] for r in episode["records"] if r["phase"] == "shared_history"],
"prior": episode["records"][0]["belief_before"], "action_costs": dict(COST),
"trained_parameters": 0, "implementation": "Lecture04-v20-LAB10"}
data = {"identity": identity, **episode}
record = json.dumps(data, ensure_ascii=False, indent=2)
(folder / "record.json").write_text(record + "\n", encoding="utf-8")
statement = f'本次运行 `{run_id}`:固定规则环境 `{truth}`、历史 `{identity["shared_history"]}`、先验 `{identity["prior"]}` 与成本 `{COST}`;本轮选择的干预是 `{intervention}`。与同条件另一轮比较时只改 `INTERVENTION`。'
result = intervention_markdown(episode)
fig, ax = plt.subplots(figsize=(10, 4.8), constrained_layout=True)
plot_weights(episode, ax, "THIS RUN | " + truth + " environment | intervention = " + intervention)
fig.savefig(folder / "candidate-weights.png", dpi=180, bbox_inches="tight")
plt.close(fig)
paths = {name: str(folder / name) for name in ("record.json", "candidate-weights.png", "result.md", "report.html")}
(folder / "result.md").write_text(statement + "\n\n" + result + "\n\n\n\n完整记录:[record.json](record.json)\n", encoding="utf-8")
# HTML is a portable saved report, generated from the same actual rows.
headers = ["阶段", "实际动作", "按钮预测", "接近预测", "实际返回", "更新后权重(按钮 / 接近)"]
rows = []
for r in episode["records"]:
cells = [r["phase"], r["action"], "开" if r["predictions"]["button"] else "关", "开" if r["predictions"]["proximity"] else "关", r["observation"]["position"] + " / " + ("开" if r["observation"]["door_open"] else "关"), " / ".join(f'{r["belief_after"][h]:g}' for h in HYPOTHESES)]
rows.append("<tr>" + "".join("<td>" + html.escape(value) + "</td>" for value in cells) + "</tr>")
summary = result.split("\n\n")[-1].replace("**", "").replace("`", "")
report = '<!doctype html><html lang="zh-CN"><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>LAB10 本次结果</title><style>body{font:17px/1.8 system-ui;margin:24px auto;padding:0 20px;max-width:1000px;color:#182b43}h1{color:#003262}img{width:100%;height:auto}.table{overflow:auto}table{border-collapse:collapse;min-width:660px;width:100%}td,th{padding:9px;border-bottom:1px solid #ccd4df;text-align:left}th{background:#eef3f8}pre{overflow:auto;white-space:pre-wrap}a{color:#003262}</style><h1>LAB10|本次干预结果</h1><p>' + html.escape(statement.replace('`','')) + '</p><div class="table"><table><thead><tr>' + ''.join('<th>'+h+'</th>' for h in headers) + '</tr></thead><tbody>' + ''.join(rows) + '</tbody></table></div><p>' + html.escape(summary) + '</p><img src="candidate-weights.png" alt="本次实际记录的候选权重;阶段、动作和开关结果来自记录"><p><a href="record.json">完整记录 record.json</a> · <a href="result.md">图表说明 result.md</a></p><details><summary>进一步检查:完整 JSON</summary><pre>' + html.escape(record) + '</pre></details></html>'
(folder / "report.html").write_text(report, encoding="utf-8")
(folder / "sha256.json").write_text(json.dumps({name: hashlib.sha256((folder/name).read_bytes()).hexdigest() for name in paths}, indent=2) + "\n")
if display_result:
from IPython.display import display, Markdown, Image, HTML
display(Markdown(statement + "\n\n" + result))
display(Image(filename=str(folder / "candidate-weights.png")))
display(Markdown("**本次保存位置:**\n\n" + "\n\n".join(f'`{path}`' for path in paths.values())))
display(HTML('<details><summary>进一步检查:完整 JSON</summary><pre>' + html.escape(record) + '</pre></details>'))
return {"run_id": run_id, "directory": str(folder), "paths": paths, "record": data}
def draw_results(data, output):
colors = {"button": "#003262", "proximity": "#c38b16"}
baseline_key = data["identity"]["baseline_episode"]
main = data["episodes"][baseline_key]
fig, axs = plt.subplots(1, 2, figsize=(12, 4.4), constrained_layout=True)
matrix = np.array([[int(predict(h, a)) for a in ACTIONS] for h in HYPOTHESES])
axs[0].imshow(matrix, cmap=matplotlib.colors.ListedColormap(["#edf1f5", "#b2d9d5"]), vmin=0, vmax=1)
axs[0].set_xticks(range(4), ["Approach\n+ press", "Press\nfrom far", "Approach\nonly", "Wait\nfar"])
axs[0].set_yticks(range(2), ["Button model", "Proximity model"])
for i in range(2):
for j in range(4):
axs[0].text(j, i, "OPEN" if matrix[i, j] else "CLOSED", ha="center", va="center", fontsize=11)
axs[0].set_title("Predictions BEFORE action", pad=18, fontweight="bold")
axs[0].set_xlabel("Same prediction Different predictions Same", labelpad=15, fontsize=9)
plot_weights(main, axs[1], "Saved baseline: " + baseline_key)
fig.suptitle("LAB10 | Baseline predictions and actual environment feedback", fontsize=14, fontweight="bold")
fig.savefig(output/"action-feedback.png", dpi=180, bbox_inches="tight")
plt.close(fig)
fig, axs = plt.subplots(1, 2, figsize=(11, 4.4), constrained_layout=True)
for ax, mechanism in zip(axs, HYPOTHESES):
row = data["memory"][mechanism]
vals=[int(r["observation"]["door_open"]) for r in row["history"]]
ax.step(range(4), vals, where="mid", lw=2.5, color=colors[mechanism])
ax.scatter(range(4), vals, s=75, color=colors[mechanism])
ax.set_xticks(range(4), ["Both", "Wait far", "Approach", "Wait far"])
ax.set_yticks([0,1], ["CLOSED", "OPEN"]);ax.set_ylim(-.45,1.35)
ax.axvspan(2.75,3.3,color="#e5e7eb")
ax.text(3, -.3, "Same NOW", ha="center",fontweight="bold",fontsize=11)
ax.set_title("Hidden rule: "+mechanism, fontweight="bold")
memory_open = row["actual_feedback"]["observation"]["door_open"]
no_memory_open = row["without_memory_actual_feedback"]["observation"]["door_open"]
ax.set_xlabel("History -> "+row["with_memory_action"]+"; cost "+str(row["with_memory_cost"])+"; observed "+("OPEN" if memory_open else "CLOSED")+"\nNo history -> "+row["without_memory_action"]+"; cost "+str(row["without_memory_cost"])+"; observed "+("OPEN" if no_memory_open else "CLOSED"),labelpad=12)
fig.suptitle("Same current observation | Both policy branches actually executed", fontsize=14, fontweight="bold")
fig.savefig(output/"memory-comparison.png", dpi=180, bbox_inches="tight")
plt.close(fig)
fig, ax = plt.subplots(figsize=(8,3.8), constrained_layout=True)
vals=[main["candidate_information_bits"][a] for a in ACTIONS]
ax.bar(["Approach + press", "Press from far", "Approach only", "Wait far"], vals,
color=["#9ca3af", "#003262", "#c38b16", "#9ca3af"])
ax.set_ylim(0,1.2);ax.set_yticks([0,.5,1]);ax.set_ylabel("Expected information (bits)")
ax.set_title("More repeated observations are not always more discriminating",fontweight="bold")
ax.text(.5,1.08,"Before the intervention; two deterministic candidates with equal prior",ha="center",transform=ax.transAxes,fontsize=9)
fig.savefig(output/"action-information.png",dpi=180,bbox_inches="tight")
plt.close(fig)
def run(output):
output=Path(output);output.mkdir(parents=True,exist_ok=True)
data={"identity": {"lab": "LAB10", "kind": "deterministic_rule_environment",
"baseline_episode": "button__approach",
"trained_parameters": 0, "seed": "not applicable; deterministic",
"goal": "Observe open door on next step, not traverse doorway",
"prior": {"button": .5,"proximity": .5},
"action_costs": COST,
"assumptions": ["mechanism fixed within episode", "no fault, noise or delay",
"button pulse and sensing evaluated afresh each step",
"two candidates do not exhaust real door mechanisms"]},
"episodes":{},"memory":{},"checks":verify_separation()}
# 真相只在构造独立环境时使用;控制器仅收到其 step 方法。
for mechanism in HYPOTHESES:
for intervention in ("approach","press_far","both"):
env=DoorEnvironment(mechanism)
result=controller_episode(env.step,intervention)
data["episodes"][mechanism+"__"+intervention]=result
assert result["goal_open_succeeded"]
env=DoorEnvironment(mechanism)
memory_result=memory_controller(env.step, keep_history=True)
no_history_env=DoorEnvironment(mechanism)
no_history_result=memory_controller(no_history_env.step, keep_history=False)
assert memory_result["history"] == no_history_result["history"]
assert memory_result["same_current_observation"] == no_history_result["same_current_observation"]
assert memory_result["actual_feedback"]["observation"]["door_open"]
assert no_history_result["actual_feedback"]["observation"]["door_open"]
memory_result["without_memory_actual_feedback"] = no_history_result["actual_feedback"]
memory_result["without_memory_decision_belief"] = no_history_result["decision_belief"]
data["memory"][mechanism]=memory_result
assert data["memory"]["button"]["same_current_observation"] == data["memory"]["proximity"]["same_current_observation"]
assert data["memory"]["button"]["with_memory_action"] != data["memory"]["proximity"]["with_memory_action"]
for mechanism in HYPOTHESES:
assert data["episodes"][mechanism+"__both"]["final_belief"] == {"button": .5,"proximity": .5}
data["checks"]["six_episode_branches_completed"]=True
data["checks"]["same_observation_different_history_actions"]=True
data["checks"]["both_memory_policies_actually_executed"]=True
(output/"trajectories.json").write_text(json.dumps(data,ensure_ascii=False,indent=2),encoding="utf-8")
draw_results(data,output)
print(json.dumps({"output":str(output),"checks":data["checks"],
"main_next_action":data["episodes"]["button__approach"]["next_action"]},ensure_ascii=False,indent=2))
return data
查看可执行代码
all_results = run("results")
查看保存的计算输出
{
"output": "results",
"checks": {
"external_feedback_controls_update": true,
"candidate_failure_is_reported": true,
"six_episode_branches_completed": true,
"same_observation_different_history_actions": true,
"both_memory_policies_actually_executed": true
},
"main_next_action": "press_far"
}
改一次干预,紧接着读本次结果
先运行默认 approach,再只把 INTERVENTION 改为 both 并运行下面这一段。每次立即得到条件说明、实际记录表、候选权重图和独立保存目录,最后可展开完整 JSON。TRUTH 只用于构造环境,控制器仍只收到 step(action)。无需先读长 JSON。
查看可执行代码
TRUTH = "button" # 两轮保持不变
INTERVENTION = "approach" # 下一轮只把这一项改成 "both"
environment = DoorEnvironment(TRUTH)
my_run = controller_episode(environment.step, INTERVENTION)
my_saved = present_intervention(my_run, TRUTH, INTERVENTION, "my_results")
本次运行 button__approach__20261001T025909.716281Z__ab311b40:固定规则环境 button、历史 ['both', 'wait_far']、先验 {'button': 0.5, 'proximity': 0.5} 与成本 {'both': 2.2, 'press_far': 1.0, 'approach': 1.2, 'wait_far': 0.0};本轮选择的干预是 approach。与同条件另一轮比较时只改 INTERVENTION。
| 阶段 | 实际动作 | 按钮候选预测 | 接近候选预测 | 环境实际返回 | 更新后权重(按钮 / 接近) |
|---|---|---|---|---|---|
| 共同历史 | both |
开 | 开 | 近 / 开 | 0.5 / 0.5 |
| 共同历史 | wait_far |
关 | 关 | 远 / 关 | 0.5 / 0.5 |
| 本次干预 | approach |
关 | 开 | 近 / 关 | 1 / 0 |
| 下一动作(已执行) | press_far |
开 | 关 | 远 / 开 | 1 / 0 |
本次干预返回关门,候选熵减少 1 bit;下一动作是 press_far,独立环境实际返回开门。完成开门目标与取得区分信息分别检查。

本次保存位置:
my_results/button__approach__20261001T025909.716281Z__ab311b40/record.json
my_results/button__approach__20261001T025909.716281Z__ab311b40/candidate-weights.png
my_results/button__approach__20261001T025909.716281Z__ab311b40/result.md
my_results/button__approach__20261001T025909.716281Z__ab311b40/report.html
进一步检查:完整 JSON
{
"identity": {
"run_id": "button__approach__20261001T025909.716281Z__ab311b40",
"generated_at_utc": "20261001T025909.716281Z",
"kind": "actual_independent_rule_environment_execution",
"truth_selected_by_experimenter": "button",
"intervention": "approach",
"shared_history": [
"both",
"wait_far"
],
"prior": {
"button": 0.5,
"proximity": 0.5
},
"action_costs": {
"both": 2.2,
"press_far": 1.0,
"approach": 1.2,
"wait_far": 0.0
},
"trained_parameters": 0,
"implementation": "Lecture04-v20-LAB10"
},
"records": [
{
"phase": "shared_history",
"action": "both",
"predictions": {
"button": true,
"proximity": true
},
"observation": {
"position": "near",
"door_open": true
},
"belief_before": {
"button": 0.5,
"proximity": 0.5
},
"belief_after": {
"button": 0.5,
"proximity": 0.5
},
"entropy_before_bits": 1.0,
"entropy_after_bits": 1.0
},
{
"phase": "shared_history",
"action": "wait_far",
"predictions": {
"button": false,
"proximity": false
},
"observation": {
"position": "far",
"door_open": false
},
"belief_before": {
"button": 0.5,
"proximity": 0.5
},
"belief_after": {
"button": 0.5,
"proximity": 0.5
},
"entropy_before_bits": 1.0,
"entropy_after_bits": 1.0
},
{
"phase": "intervention",
"action": "approach",
"predictions": {
"button": false,
"proximity": true
},
"observation": {
"position": "near",
"door_open": false
},
"belief_before": {
"button": 0.5,
"proximity": 0.5
},
"belief_after": {
"button": 1.0,
"proximity": 0.0
},
"entropy_before_bits": 1.0,
"entropy_after_bits": -0.0
},
{
"phase": "next_goal_action",
"action": "press_far",
"predictions": {
"button": true,
"proximity": false
},
"observation": {
"position": "far",
"door_open": true
},
"belief_before": {
"button": 1.0,
"proximity": 0.0
},
"belief_after": {
"button": 1.0,
"proximity": 0.0
},
"entropy_before_bits": -0.0,
"entropy_after_bits": -0.0
}
],
"candidate_information_bits": {
"both": 0.0,
"press_far": 1.0,
"approach": 1.0,
"wait_far": 0.0
},
"final_belief": {
"button": 1.0,
"proximity": 0.0
},
"next_action": "press_far",
"goal_open_succeeded": true
}独立阅读材料:原基线、候选信息量与记忆对照
下面三张图读取 all_results = run("results") 的六情节与记忆对照。第一张明确对应按钮环境+approach 的基线;第二张对应等权先验下的四动作候选信息量;第三张对应两种隐藏机制下、有历史与无历史策略的实际执行。它们不随上方单次干预改变。本次干预图已在上方出现,并保存在 my_saved["directory"]。
查看可执行代码
fig, ax = plt.subplots(figsize=(12, 4.8))
ax.imshow(plt.imread("results/action-feedback.png"))
ax.axis("off")
plt.tight_layout()
plt.show()

查看可执行代码
fig, ax = plt.subplots(figsize=(12, 4.8))
ax.imshow(plt.imread("results/action-information.png"))
ax.axis("off")
plt.tight_layout()
plt.show()

查看可执行代码
fig, ax = plt.subplots(figsize=(12, 4.8))
ax.imshow(plt.imread("results/memory-comparison.png"))
ax.axis("off")
plt.tight_layout()
plt.show()
