如何高效使用 Codex
一份面向实际开发、科研、机器学习与多 Agent 协作的实战使用手册。
0. 核心结论
高效使用 Codex,最重要的转变是:把 Codex 当作一个能够读取代码库、执行命令、修改文件、运行测试、检查结果并继续迭代的软件工程 Agent。
你的主要工作逐渐从“告诉它每一行代码怎么写”,转变为“定义目标、提供必要上下文、设置边界、规定验收标准,然后检查结果”。
一个高质量 Codex 任务通常应该包含四件事:
目标
上下文
约束
验收标准
例如,与其说:
帮我改一下 train.py。
更好的说法是:
修复 train.py 中 resume training 后 learning-rate scheduler 状态没有正确恢复的问题。
背景:
- checkpoint 已保存 optimizer 和 scheduler state
- 当前 resume 后 loss 曲线会突然异常
- 不希望改变正常从头训练的行为
要求:
1. 先定位问题,不要直接猜。
2. 尽量做最小修改。
3. 增加一个可以复现问题的测试。
4. 运行相关测试。
5. 最后告诉我:
- 根因
- 修改了哪些文件
- 如何验证
- 是否还存在风险
1. 先理解 Codex 的工作方式
Codex 最有效的工作方式大致是:
理解目标
↓
读取相关代码
↓
形成假设 / 计划
↓
修改代码
↓
运行测试 / build / benchmark
↓
查看结果
↓
修复问题
↓
再次验证
↓
完成
长任务尤其依赖这个循环。真正决定产出的,往往不是 prompt 写得多复杂,而是 Codex 是否能进入一个稳定的“行动—反馈—修正”闭环。
2. 选择正确的 Codex 使用入口
2.1 Desktop / GUI
适合日常开发、查看 diff、多任务并行、Worktree 管理、长任务以及不想一直盯着 Terminal 的工作。
2.2 CLI
codex
适合 SSH、远程服务器、GPU 训练机、Linux 环境、快速检查代码、直接跑测试和日志分析。
2.3 IDE Extension
适合一边人工写代码一边使用 Codex,尤其适合精确查看函数、小范围 refactor、debugging 和代码理解。
2.4 Codex Cloud
适合本机不想被长期占用、任务可能运行较久、希望电脑休眠后任务仍可继续的场景。
Local Codex
→ 在你的机器上执行
Codex Cloud
→ 在托管环境中执行
2.5 Remote
如果开发机位于宿舍、办公室或服务器旁边,而你手上只有另一台设备,可以把 Remote 理解为远程控制平面:代码仍运行在开发机上,你负责启动、审批、steering、review diff 和查看结果。
3. 仓库初始化
第一次让 Codex 接触一个项目时,不要马上扔给它复杂任务。先确认 Git 状态与项目运行方式。
git status
推荐第一条任务:
先不要修改代码。
请阅读这个仓库,告诉我:
1. 项目的主要目录结构。
2. 主程序入口。
3. 测试方式。
4. build / lint / typecheck 命令。
5. 哪些文件最重要。
6. 如果以后让你开发功能,你认为还缺哪些项目说明。
不要逐文件复述。
4. AGENTS.md
AGENTS.md 是 Codex 项目里最重要的长期上下文文件之一,适合保存项目结构、常用命令、约束和编码规范。
项目结构
## Architecture
- `src/model/`: model definitions
- `src/train/`: training loop
- `src/eval/`: evaluation
- `scripts/`: experiment scripts
常用命令
## Commands
Install:
pip install -e .
Tests:
pytest tests/
Lint:
ruff check .
Type check:
pyright
约束
## Constraints
- Do not change dataset preprocessing unless explicitly requested.
- Preserve checkpoint backward compatibility.
- Do not commit generated checkpoints.
5. 不要把 AGENTS.md 写成百科全书
长期积累的大量旧规则很容易产生反作用。
不推荐:
每次修改代码之前:
阅读 architecture.md
阅读 database.md
阅读 deployment.md
阅读 README.md
读取全部 tests
更合理:
Use architecture.md when changing service boundaries.
Use database.md when changing schemas.
Use deployment.md for deployment-related work.
6. 如何写一个高质量 Codex Prompt
任务
背景
约束
验收
输出要求
任务:
<你希望最终发生什么>
背景:
<为什么要做>
<相关模块>
<已有情况>
约束:
- 不要修改……
- 保持……
- 尽量……
- 可以修改……
验收:
- 测试 X 通过
- benchmark 达到 X
- 行为 X 保持不变
- case X 可以正确工作
完成后:
1. 总结根因/设计。
2. 列出修改文件。
3. 告诉我运行了哪些验证。
4. 报告仍然存在的风险。
7. 少告诉 Codex“怎么做”,多告诉它“什么算完成”
低效:
打开 a.py。
找到 foo。
增加一个 if。
然后修改 b.py。
更好:
解决用户 logout 后 websocket 连接仍然存在的问题。
要求:
- logout 后 1 秒以内关闭连接
- 正常断线重连逻辑不能受影响
- 添加 regression test
先调查目前 websocket lifecycle,再选择最合理的实现。
8. 什么时候使用 Plan
复杂任务可以先让 Codex 调查而不修改:
先调查,不要修改代码。
请:
1. 找到相关实现。
2. 描述当前数据流。
3. 找出最可能的问题。
4. 提出修改方案。
5. 说明涉及哪些文件。
6. 说明测试方案。
等方案明确后再实施。
Plan 特别适合架构变化、数据库 migration、大范围 refactor、陌生代码库、高风险功能和跨多个 subsystem 的任务。
9. Goals:长任务的关键工具
Goal 更适合“路径未知,但终点明确”的任务,例如性能优化、flaky test、benchmark 优化、debugging、大型 refactor 和多轮实验。
/goal Reduce p95 inference latency below 100 ms while keeping all correctness tests green.
机器学习任务尤其适合:
/goal Reduce 4-NFE PPL below 350 on the fixed validation configuration,
while preserving entropy and distinct-2 within the baseline range.
Use the existing evaluation scripts.
Do not change the evaluation dataset.
Record each experiment and its result.
一个优秀 Goal 通常包含:
Outcome
Verification
Constraints
Boundaries
Iteration policy
Stop condition
10. Worktree:并行开发最重要的工具之一
main
├── worktree A
├── worktree B
└── worktree C
适合 Worktree 的场景包括:多个 Codex 并行、Codex + Claude Code、实验性 refactor、长任务以及不确定是否最终采用的方案。
11. 如何并行使用多个 Agent
不要让多个 Agent 同时改同一个核心文件。更好的并行方式是按问题或模块划分。
Agent A → 调查训练速度瓶颈
Agent B → 检查 dataloader
Agent C → 分析 profiler 输出
Agent D → 检查 GPU utilization
或者:
Agent A → frontend
Agent B → backend
Agent C → tests
多 Agent 最适合:独立调查、独立模块、独立实验、独立 review。
12. Codex 最擅长的任务
12.1 Bug investigation
Investigate this bug.
Observed:
<现象>
Expected:
<预期>
Reproduction:
<复现方法>
Do not patch immediately.
First:
1. reproduce it;
2. trace the relevant execution path;
3. identify the root cause.
Then implement the smallest robust fix and add a regression test.
12.2 新功能实现
Implement <feature>.
Before editing:
- inspect the existing architecture;
- reuse existing abstractions when appropriate;
- identify the smallest integration point.
Requirements:
...
Acceptance criteria:
...
Run relevant tests after implementation.
12.3 Refactor
Refactor the evaluation pipeline to separate:
- generation
- metric computation
- result persistence
Do not change observable evaluation behavior.
After refactoring:
- existing tests must pass;
- add tests if current coverage is insufficient;
- compare outputs on a small fixed fixture.
12.4 读代码
Trace what happens from:
python train.py ...
until the first optimizer.step().
Focus only on:
- config loading
- dataset creation
- model creation
- loss computation
- backward
- optimizer update
Give me file/function references.
13. 让 Codex 使用真实证据
重要任务中应明确要求:
Do not consider the task complete until the relevant verification has been run.
验收证据优先级:
真实运行结果
>
自动化测试
>
静态检查
>
代码推理
>
“看起来应该可以”
14. 让测试成为 Definition of Done
完成条件:
pytest tests/test_checkpoint_resume.py
必须通过。
另外运行一个 100-step smoke training,
确认 resume 前后的 LR 连续。
一旦完成条件明确,Agent 的搜索空间会明显缩小。
15. 长任务一定要外部化状态
长任务不要把所有状态只留在聊天记录里。建议把关键状态写入 repo:
docs/
architecture.md
experiments/
README.md
results.csv
PLAN.md
STATUS.md
特别长的任务可以维护:
# Goal
# Current understanding
# Completed
# Current hypothesis
# Failed approaches
# Next step
# Verification
16. Side Chat
当主线程正在执行长任务,而你只是想问一个解释性问题时,Side Chat 可以避免主任务被打断。
/side
适合:问解释、问架构、问错误、讨论想法、分析某个 diff。
17. 什么时候开新 Chat
同一个目标:继续原 chat。
新的独立目标:开新 chat。
原因是旧线程会积累大量与新任务无关的上下文。
18. 上下文越来越长怎么办
/compact
也可以主动要求 Codex 把当前状态写入 STATUS.md,然后新线程从文件继续。
19. 使用 /review
Review the current diff as if you were reviewing a pull request.
Focus on:
- correctness
- regressions
- edge cases
- concurrency
- state management
- performance
- test coverage
Do not modify anything yet.
Rank findings by severity.
一个非常有效的结构:
Agent A 写代码
↓
Agent B review
↓
Agent A 修复
20. Git 工作流
git status
git diff
git log
推荐流程:
main
↓
create branch/worktree
↓
Codex implementation
↓
tests
↓
review diff
↓
commit
↓
merge / PR
大型 refactor、实验代码、依赖升级、migration 等任务,不建议直接让 Agent 在 main 上无限修改。
21. 权限和安全
本地 Codex 可以执行 shell command,因此权限设计很重要。
普通任务通常只需要允许:
读取项目
修改项目
运行测试
以下操作应谨慎:
删除大量文件
修改系统配置
访问 secrets
push production
数据库 destructive migration
sudo
22. Secret 管理
不要把 API key、密码、access token、private key 直接写入 AGENTS.md、prompt、源码或 commit。
优先使用:
environment variables
secret manager
.env(并加入 .gitignore)
Agent 通常只需要知道 secret 从哪里读取,而不需要知道 secret 本身。
23. Skills
如果经常重复同一种复杂流程,可以把它做成 Skill,例如:
run-training-evaluation
prepare-release
review-migration
benchmark-model
generate-experiment-report
可以这样理解:
AGENTS.md
= 项目宪法
Skill
= 标准作业流程
Prompt
= 当前工单
24. MCP / Plugins
当 Codex 需要频繁访问外部系统时,可以接入 GitHub、documentation、issue tracker、database、internal services 等。
这通常比把大量网页内容复制到 prompt 中更适合长期项目。
25. 模型如何选择
简单任务:改变量名、增加 log、修 typo、简单测试。
中型任务:普通功能、debug、模块重构。
困难任务:架构、复杂 race condition、大型 migration、科研代码、复杂数学逻辑、长任务。
原则是:不要把大量时间花在“哪个模型理论上聪明 3%”,而应把更强模型留给真正需要复杂推理和长程规划的任务。
26. 减少 Token 浪费
- 不要每次都让 Codex 读取整个 repo。
- 长期背景放 AGENTS.md / docs,不要反复塞进 prompt。
- 不要不断发送“继续”;如果终点明确,考虑 Goal。
- 简单任务不要要求长篇解释,只保留 changed files / verification / remaining risks。
27. 最容易失败的 Codex 使用方式
27.1 任务过于模糊
优化一下代码。
更好:
Reduce peak GPU memory during evaluation by at least 20%
without changing evaluation outputs.
27.2 同时塞十个无关任务
UI、API、PyTorch 升级、CUDA 优化、数据库重构、README 最好拆成独立任务。
27.3 没有 Definition of Done
没有验收标准时,Agent 很容易停在“看起来合理”。
27.4 用户自己规定错误实现
如果你不确定实现方式,不要强制规定某个技术方案;可以表达假设,让 Codex 先验证。
27.5 没有先复现 bug
28. 科研 / Machine Learning 项目的推荐结构
project/
├── src/
├── configs/
├── scripts/
├── tests/
├── experiments/
│ ├── README.md
│ ├── results.csv
│ └── notes/
├── docs/
│ └── architecture.md
├── AGENTS.md
└── README.md
AGENTS.md 可加入实验约束:
## Experiment policy
- Do not overwrite existing experiment results.
- Every new experiment must record its config.
- Use fixed evaluation seeds unless explicitly testing variance.
- Never change evaluation protocol when comparing experiments.
- Record failed experiments as well as successful ones.
29. 让 Codex 做实验
Goal:
Improve 4-step generation quality.
Success criteria:
- PPL < 350
- entropy remains within baseline ±5%
- distinct-2 does not decrease more than 5%
Constraints:
- same validation set
- same tokenizer
- same evaluation script
- no cherry-picking samples
Workflow:
1. inspect current implementation;
2. propose one hypothesis at a time;
3. make the smallest testable change;
4. run the defined experiment;
5. record result;
6. keep or revert based on evidence.
Stop if three consecutive hypotheses fail to improve the primary metric and report findings.
30. Codex + ChatGPT 的推荐分工
ChatGPT:理论分析、架构讨论、阅读论文、设计实验、比较方案。
Codex:读取 repo、实现、debug、测试、benchmark、Git、真实运行。
ChatGPT
↓
确定方向
↓
SPEC.md
↓
Codex
↓
实现 + 测试
↓
实验结果
↓
ChatGPT
↓
分析结果
31. Codex + Claude Code
不推荐让 Codex 和 Claude Code 同时编辑同一个 worktree。
更好的方式:
worktree-codex
worktree-claude
或者:
Codex → implementation
Claude → review
再或者:
Claude → architecture investigation
Codex → implementation
32. 一个完整的标准工作流
- 创建 Worktree:一个任务一个 worktree。
- 让 Codex 调查:
Investigate first. Do not edit yet. - 确认目标、约束、验收。
- 复杂任务先 Plan。
- 执行 implementation。
- 运行 tests / lint / typecheck / benchmark。
- 使用 review。
- 修复 review 问题。
- 查看
git diff。 - Commit。
33. 常用 Prompt 模板库
模板 1:理解项目
Do not modify anything yet.
Inspect this repository and explain:
1. high-level architecture;
2. major execution entry points;
3. important modules;
4. test/build/lint workflow;
5. where configuration lives;
6. areas that appear fragile or unusually complex.
Do not summarize every file.
Focus on information useful for future development.
模板 2:Bug
Investigate the following bug:
Observed:
...
Expected:
...
Reproduction:
...
First reproduce and identify the root cause.
Do not patch based only on speculation.
Then:
1. implement the smallest robust fix;
2. add a regression test;
3. run relevant tests;
4. report root cause, changes and remaining risks.
模板 3:Feature
Implement:
...
Requirements:
...
Constraints:
...
Acceptance criteria:
...
Before editing, inspect the existing architecture and reuse existing abstractions where appropriate.
After implementation, run the relevant tests and report the verification results.
模板 4:Refactor
Refactor:
...
Goal:
...
The observable behavior must remain unchanged.
Before editing:
- identify current interfaces;
- identify dependencies;
- identify tests covering the behavior.
After editing:
- run regression tests;
- compare behavior before/after where practical;
- report any assumptions.
模板 5:Performance
Goal:
Improve <metric> from X to Y.
Constraints:
- correctness must remain unchanged;
- do not change benchmark methodology;
- do not reduce workload;
- do not disable safety checks.
Workflow:
1. establish baseline;
2. profile;
3. identify bottleneck;
4. make one meaningful optimization;
5. rerun benchmark;
6. compare results.
Continue until the goal is reached or no defensible optimization remains.
模板 6:代码 Review
Review the current diff.
Do not modify files.
Look specifically for:
- correctness bugs
- regressions
- edge cases
- state bugs
- concurrency issues
- performance regressions
- security issues
- inadequate tests
Prioritize findings by severity.
Ignore purely stylistic comments unless they materially affect maintainability.
34. 一个高质量任务应具备的信息
- Codex 知道我要实现什么吗?
- Codex 知道为什么要实现吗?
- Codex 知道什么不能改变吗?
- Codex 知道如何判断完成吗?
- Codex 可以运行验证吗?
- 这是一个独立任务吗?
- 是否应该创建 worktree?
- 是否复杂到应该先 plan?
- 是否长到应该使用 goal?
35. 最终原则
- 给目标,不要过度微操实现。
- 明确 Definition of Done。
- Bug 先复现,再修。
- 重要修改必须运行真实验证。
- 一个独立任务一个 Chat / Worktree。
- 复杂任务先 Plan。
- 路径未知、终点明确的长任务使用 Goal。
- 长期信息放 AGENTS.md / docs,不要重复塞进 prompt。
- AGENTS.md 保持精简,避免积累过时规则。
- 把 Codex 的输出当作需要 review 的工程成果,而不是天然正确的答案。
一句话工作流
定义结果
→ 给必要上下文
→ 设边界
→ 定义验收标准
→ Codex 调查
→ 实现
→ 实际运行
→ Review
→ Commit