
- 本《Codex零基础入门教程》保证你和你的兄弟姐妹都能看懂,只要小学毕业就能看明白。这不是功能词典,而是一名普通学习者从第一次打开 Codex,到真正让它完成一个可验收任务的完整路线,以最通俗易懂的语言传递给你。建议您先收藏,后边慢慢看!
第一次打开 Codex 时,我犯的错误和很多人一样:把它当成一个“更会写代码的 ChatGPT”。我在输入框里写了一句“帮我做一个网站”,然后盯着它生成文件。几分钟后页面确实能打开,但按钮有的不能点,移动端会溢出,刷新后数据消失,我也不知道它改了哪些文件。
那一刻我才明白:会生成代码,不等于会交付项目。Codex 真正厉害的地方,不是某一次回答写得多漂亮,而是它能进入一个真实项目,读取文件、理解约束、修改代码、运行命令、检查结果,再根据反馈继续迭代。它更像一个能操作电脑和项目环境的执行者,而不是只在聊天框里给建议的问答机器人。
但这也意味着,使用 Codex 的门槛并不只是“会不会写提示词”。你还要学会给它正确的工作目录、合适的权限、足够但不过量的上下文,以及明确的验收标准。只要这四件事没处理好,再强的模型也可能在错误方向上跑得很快。
这篇教程不要求你是程序员。你只需要会创建文件夹、安装软件、复制命令,并愿意在每一步查看结果。我会用一个“个人任务看板”作为练习项目,把安装、第一次对话、需求拆解、修改文件、运行测试、代码审查、长期规则和重复工作复用全部串起来。跟着做完,你得到的不只是一个 Demo,而是一套以后做网页、自动化脚本、数据工具和 AI 产品都能复用的方法。
一、先理解 Codex:它不是“替你写几段代码”,而是替你完成一段工作循环

过去我使用普通 AI 编程工具时,流程通常是:我描述问题,AI 给出代码,我复制到编辑器,报错后再把报错复制回来。上下文在聊天框、编辑器和终端之间来回搬运,最累的不是写代码,而是不断解释“我刚才做了什么”。
Codex 把这条链路接了起来。它可以在授权范围内查看项目文件、搜索代码、编辑文件、执行构建或测试命令,并把结果继续作为下一步判断依据。一个完整循环通常是:
- 理解目标:确认你到底要解决什么问题。
- 检查环境:读取目录、关键文件、项目说明和依赖。
- 制定计划:把大任务拆成可验证的小步骤。
- 执行修改:创建或编辑文件,必要时运行命令。
- 验证结果:执行测试、构建、类型检查或实际预览。
- 审查差异:确认没有误改文件,没有引入明显回归。
- 交付说明:告诉你改了什么、验证了什么、还剩什么风险。
真正有价值的是第五步和第六步。只生成代码的 Agent 很容易给你一种“已经完成”的错觉,而一个合格的 Agent 应该拿证据证明结果。以后每次下任务,我都会在结尾加一句:
完成后请运行与本次修改相关的检查,并告诉我:改了哪些文件、运行了哪些验证、结果如何、还有哪些未验证风险。不要只说“已完成”。
这一句看起来普通,却能明显减少“代码写完了但不能用”的情况。
二、四种入口怎么选:新手先选离工作最近的那个

Codex 目前不是单一形态。你会看到桌面应用、IDE 扩展、CLI 命令行和 Web/云端。它们不是谁替代谁,而是适合不同的工作位置。
桌面应用:适合第一次接触 Codex 的人。你可以选择本地项目、查看文件变化、开多个任务,也不必先熟悉终端。它更像一个“AI 工作台”,不仅能处理代码,也能处理文档、表格、网页和其他文件型任务。
IDE 扩展:适合已经在 VS Code、Cursor 或 Windsurf 里写代码的人。它能直接利用当前打开的文件、选中的代码和编辑器上下文。小范围修改、解释代码、修复当前报错时,IDE 入口通常最顺手。
CLI:适合想把 Codex 放进终端工作流的人。它能在项目目录中直接运行,适合批量改动、脚本化、CI 或需要频繁执行命令的任务。本文会重点讲 CLI,因为它最能让你看清 Codex 如何读取项目、申请权限和验证结果。
Web/云端:适合把任务交给远程环境长时间运行,或者并行处理多个项目问题。它的优势不是“界面更简单”,而是任务可以脱离你当前电脑持续执行。但云端环境与本地环境并不完全相同,依赖、密钥、网络权限需要单独配置。
我的选择方法很简单:第一次学习用桌面应用或 CLI;正在写代码时用 IDE;需要并行或长时间执行时再用云端。不要一开始把四种入口全部配置一遍。入口越多,不代表效率越高,反而容易把注意力耗在配置上。
三、开始前只准备三样东西:项目文件夹、Git 和一个可验证的小目标
很多教程一上来就讲模型、MCP、Skills 和复杂配置。我照着配置了一堆东西,最后连第一个任务都没跑通。后来我把准备工作缩成三项。
第一,准备一个独立项目文件夹。不要第一次就让 Codex 扫描桌面、下载目录或整个硬盘。工作目录既决定它看到什么,也决定它默认能改什么。新手最好创建一个专门练习目录:
mkdir codex-first-project
cd codex-first-project
第二,安装 Git,并养成任务前后留检查点的习惯。Git 不是程序员专属工具,它更像项目的“撤销历史”。在 Codex 动手前提交一次,任务完成后再看差异,即使修改不满意,也能准确知道发生了什么。
git init
git add .
git commit -m “before codex task”
如果目录还是空的,可以先创建一个 `README.md`,再做第一次提交。你也可以用图形化 Git 工具,不必强迫自己背命令。核心不是命令,而是给每次自动修改留下可恢复的边界。
第三,选择一个能在 30 到 60 分钟内验收的小目标。第一次不要做“完整电商平台”“微信替代品”或“全自动赚钱系统”。目标越大,你越难判断问题来自需求、模型、环境还是代码。本文的练习目标是:
做一个本地运行的个人任务看板:能新增任务、标记完成、删除任务;刷新页面后数据仍保留;界面适配手机;不接后端,不需要登录。
这个目标足够小,却包含 UI、交互、数据保存和响应式布局,正好能练习一条完整交付链路。
四、安装 Codex 和 CLI:先用官方方式,再处理系统差异

官方当前为 macOS 和 Linux 提供独立安装脚本。在终端中运行:
curl -fsSL https://chatgpt.com/codex/install.sh | sh
安装后执行:
codex
第一次启动时,按界面提示选择“使用 ChatGPT 登录”或其他可用的登录方式。登录完成后,先检查版本和帮助信息:
codex –version
codex –help
If you already have a Node.js environment, you can also choose npm from the official installed page; if you are used to Homebrew, you can also choose Homebrew from the page. Do not install it in three ways at the same time, otherwise there may be multiple `codex ' implementable documents in the system and the old version is still being updated。
Windows users give priority to the official Windows desktop application or the Windows installation path currently available on the official page. WSL can also be used if your development project is mainly run in Linux environment. The key is not “what is more advanced”, but Codex, Git, Node/Python and the project document must be in the same accessible environment. The most common Windows pit is where the project is on the Windows Disk, dependent on being installed in the WSL, and the terminal is activated from another environment, lastly as an order cannot be found or as an error of authority。
Do not start any directory directly after installation. Let's go to the practice directory
cd codex-first-project
codex
The Codex startup interface displays the current model, work catalogue and context status. Confirm that the directory is correct. If the directory is wrong, exit immediately and restart in the correct directory. Making Agent work in the wrong directory is far less dangerous than a hint。
For the first time, don't let it write the code: let it explain what it sees
For the first time, I really built trust, not because Codex created a page, but because I made it do a read-only check first. I entered:
Do not create or modify any files first. Please check the current directory and tell me:
1. What documents are available
2. What is the status of this project
What are the minimum technical options you propose to use in order to fulfil the “local personal task board”
4. What options do I need to identify。
This step has three effects. First, make sure it's really in the right directory. Secondly, to confirm that it understands the goal. And third, let me see the technical choice before writing the code. If it comes up with advice on databases, account systems, Docker and micro-services, I'll know that the program is over-designed。
An ideal answer should be to press the program to a minimum, for example, using Vite + React, or simply using HTML, CSS, JavaScript and `localStorage ' . For the first exercise, I would prefer to have three original sets because of the low level of reliance, the simplicity of start-up and the understanding of each document。
Then it turns the target into a receiving and inspection list:
Change the demand to a verifiable acceptance list. Each must be checked by page operation or command and not write unjudgeable descriptions such as “exhausted” “code grace”. Just get out of the list and don't start developing。
you should get a similar result: the page was opened without error for the first time; the task could be entered and added; the empty content could not be submitted; the completed status could be changed; the task could be deleted; the refreshed data retained; no horizontal spill at 375 px width; and there were no unprocessed control table errors。
It's only then that the demand goes from "I want to do something" to "what's done."。
VI. Not to write in papers, but to complete four fields

I later found that the main problem with the novices was not that they were not long enough, but rather that they were missing fields. Official best practices also summarize high-quality tasks in four parts: objectives, context, constraints and completion criteria。
Target(Goal): Answer “What to change”. Instead of saying “optimize”, it would say “support the submission of return vehicles when new tasks are added and prevent pure space content”。
**Context(Context): Answer “Where should we look first”. You can specify a file, directory, screenshot, error reporting, interface document or reference realization. The more context the better, the more important thing is to have Cordex read the material。
**Constraints(Constraints): The answer is “Nothing to destroy”. For example, no new framework is introduced, no public interfaces are modified, no `.env ' is read, current visual style is maintained and only current modules are modified。
**Completion standards(Done when): Answer “how to prove it”. For example, test pass, build success, specify interactive recurrences, screenshots are consistent with reference diagrams, no Lint errors are added。
Put the four parts together, and this exercise can say:
Objective: To complete a local personal task board in the current empty directory, to support the addition, completion/cancellation, deletion and permanence of the localStorage。
Context: Read README.md first; the project currently has no code. Prioritize primary HTML, CSS, JavaScript to ensure that I can understand it。
Constraint: No backend, no login, no frontend frame, no remote CDN; interfaces keep white bottom, blue emphasis; no horizontal scrolling at 375 px width of the moving end。
Completion standards: Local start-up and validation after creation of necessary documents; item-by-case check for new, empty-value intercepts, state-to-state switching, deletion, refreshing and moving-end layouts; final list of modified documents, validation steps and uncovered risks. A maximum of 6 steps plan was given before the start, and it was confirmed that the programme had not exceeded requirements before implementation。
Such hints are not fancy, but they are more useful than “You are a world-class insular engineer, please think in depth”. Role-setting cannot be a substitute for project facts, nor can emotional emphasis be a substitute for acceptance standards。
VII. APPROACH OF THE APPROACH: IT'S NOT THE BIG BUSINESS, IT'S THE COURSE

For the first time, when I saw a power hint, my instinct was to allow it in its entirety, lest it be repeated. And then I realized that Codex was able to execute orders, modify documents and access the network, and that excessive access not only increased risks, but also deprived you of the opportunity to observe its work。
Current Codex permissions can be understood as three:
- `:read-only`: Fit to read codes, explain projects, do programmes, check risks. It cannot modify the working area directly。
- `:workspace`: Allows for writing in the current working area and is limited to the selection and suitable for most day-to-day development tasks。
- `_other organiser`: Remove local sandbox restrictions. Only when you clearly need extensive access and fully understand the implications。
It is enough for the newcomer to default on the selection of working-level privileges. Read-only when exploring strange warehouses; confirm the program and then cut to the work area for writing. Do not give full permission for less than two confirmations and do not direct the directory containing private keys, wallets, production vouchers or a large number of personal documents as a workspace。
There are two different concepts:Sandbox BorderIt's not like it's the same thing. To decide where to order access,”Approving Policy”Decide when to suspend and ask you. Even in sandboxes in the work area, a mission may request authorization because of the need to network, to write an off-site directory or to perform high-risk operations。
My simple rule is to read the code only, to change the workspace for normal items, to install a dependency and access network to be confirmed on a step-by-step basis, and to remove files, reset history, enforce delivery and perform a separate check of the production environment. The purpose of the permission is not to stop the Codex job, but to make the impact radius of the error manageable。
VIII. Orders that must be known: 8 first, and other if necessary
Cordex's slash command will change with the entry and version you use. Enter `/ ' commands that can see what the current environment really supports. The new guy doesn't need a complete back list
`/status ' : view the current session, the context and the restricted status. the longer the task, the more it should be seen once in a while。
`/model ' : select a model for the current task. do not superstitiously always use the most heavy model; clear minor modifications can give priority to speed, complex structures, difficult debugging and long missions that increase the strength of reasoning。
`/ressoning ' : adjustment of reasoning input. low intensity is appropriate for well-defined tasks, medium and high intensity is appropriate for cross-document modification and debugging, and the highest is better for truly difficult and deserving tasks。
`/missions ': select the range of operations permitted by the current task. the different entry points may vary slightly, depending on the actual interface options。
`/init ': Generate an initial template `AGENTS.md ' for the current project. It is not completed after its creation and must be adapted to the actual circumstances of the project。
`/plan ' : to enter the planning model, tailoring to needs is not clear, or to set a course when the mandate has multiple phases。
`/review ' : code review for failure to submit changes, designated submissions or baseline branches. its value is to replace the “review perspective” of the examination, not to repeat what has just been achieved。
`/compact ' : compress early context when talking long, keep targets, constraints, decision-making and progress, and reduce interference with old information。
If you are in the CLI, you often resume sessions, fork sessions, etc. Official best practice refers to orders such as `/resume ' , `/fork ' , but the available orders will vary depending on the client and the environment, so the most reliable practice remains to enter `/ ' or to view the official current order page。
IX. `AGENTS.md ' : Translating repeated words into project rules

I understood the value of `AGENTS.md ' when I reminded Codex, “Don't use npm, use pnpm” “Modified Run Test” “Don't touch the Generating Directory”. It is equivalent to a project description written to Agent, and Codex automatically reads the rules at the relevant level when entering the project。
you can run `/init ' drafts in the root directory and then compress them into really useful content. a suitable version of the academy project could be:
# AGENTS.md
## PROJECT OBJECTIVES
– This is a zero-dependent local task board exercise。
– Prefer to keep the code simple, readable, and avoid excessive abstraction。
## MODIFICATION RULES
– Use original HTML, CSS, JavaScript without new frames or remote dependence。
– Do not modify files unrelated to the current task。
– Do not read or output any key, environmental variables or personal files。
## CERTIFICATION RULES
– When JavaScript is modified, check that no errors were added to the browser console。
– Validation of new, completed switching, deletion and refreshing per delivery。
- The final statement must set out the changes to the document, the validation results and the remaining risks。
`AGENTS.md ' is the easiest to make two mistakes. The first is a list of wishes written in thousands of words, filled with vague requirements such as “quality” “best practices” “deep thinking”. Too long a rule would take its context and could conflict with one another. The second is to include one-off requirements, such as “to make the button green today”. Long-term rules should be stable, enforceable and repeated。
Larger warehouses can be layered with rules: global file-keeping personal habits, common norms for the warehouse root catalogue preservation team, subdirectories with local service or modular rules. The closer to the specific rule of the current document, the higher the priority. You don't need to design a complete system on the first day; when the same error happens the second time, you add the corresponding rule。
X. REAL RUNNATION: PLAN, IMPLEMENTATION, VERIFICATION, REVIEW

Now go back to the mission board. Read-only inspections, acceptance lists and project rules have been completed, followed by Codex。
In the first round, there was no simultaneous pursuit of functionality and vision. Let's get it done with the minimum ring:
The first edition is now on the confirmed acceptance list. Availability completed without additional animation and complex components. Inspections that can be implemented at each stage of completion are run; if the environment lacks the necessary tools, the minimum solution is indicated and no unauthorized increase in dependency is allowed. Stop after completion and wait for my acceptance。
Codex should create the necessary documents `index.html ' , `styles.css ' , `app.js ' and initiate or inspect them in a manner appropriate to the current environment. You need to see which orders it actually runs. Don't just look at the final response, because the execution log tells you if it's really validated。
Once the first edition comes out, I will do a round of manual check-ups myself: three new tasks; enter spaces to see if they are stopped; finish one mark; delete the other; refresh the page to see if the status is maintained; zoom in the browser width to the cell phone size; open the console to see if it is reported。
When problems are identified, do not rub the whole phenomenon in the words “is not working”. Give Codex a minimum recurrence:
I FOUND A REPLICABLE PROBLEM: THE NEW TASK “A”, MARKED AS COMPLETED, CONTINUED AFTER THE PAGE WAS REFRESHED, BUT THE COMPLETION STATUS WAS LOST. PLEASE LOCATE THE ROOT CAUSE AND EXPLAIN WHICH STEP THE DATA IS NOT SAVED BEFORE MAKING THE MINIMUM MODIFICATIONS. DO NOT RECREATE IRRELEVANT CODES. THIS SET OF STEPS WAS RETESTED AFTER THE REPAIRS AND A CHECK WAS COMPLETED TO PREVENT THE RETURN OF THE SAME TYPE。
this formulation makes clear the phenomenon, the re-emergence steps, the expected behaviour, the modification of the boundary and the means of verification. it is easier to achieve stable results than “upload the bug, fix it”。
Once the function is completed, the visual wheel is done separately:
optimizing the interface level without changing the current functionality and data structure: white background, dark grey text, blue emphasis; maximum width of content area 720px; no spill in input area and task card below 375px width; clear but not diluted completion status. changes only styles and necessary semantic labels. compares the changes before and after completion and again validates all functions。
Finally run `/review ' , allowing Codex to check from a review point of view that no changes were submitted. The focus of the review can be on the reliability of data sustainability, the security of user input, the existence of inoperable interactions, the spillover of the mobile end, and the availability of irrelevant changes. The review found that the problem was not entirely mechanical, but whether it was real or not。
The key to this link is not “Codex right at a time”, but to turn failure into a problem of locatorability, repairability and reversion. What you really train is delivery capacity, not card draw。
XI. Context management: Too little guess, too much gets lost
The quality of Codex's output depends heavily on what it sees now. The two extremes common to newcomers are to give a single message of need and to expect it to find everything by itself, or to stuff dozens of documents, all chat records and the whole demand library at once。
I now divide the context into four layers。
- The first level is the current task: what this side does, how the problem is repeated and what the criteria are for its completion. It should be short and clear。
- The second level is the project rules: placed in `AGENTS.md ' , including the catalogue structure, operating orders, project constraints and certification requirements。
- The third level is reusable: when the same set of steps is repeated, it is sealed with Skill. For example, a fixed code review list, issuance of notes generation, log referencing process, content layout specifications。
- The fourth level is external dynamic information: links through MCP or plugins when data are available in GitHub, Setry, Figma, Notion, database or other systems and are subject to continuous change. Do not access all tools for “looks professional”; each additional tool adds a set of privileges, methods of failure and context noise。
When the session gets longer, let Codex sort it out and then use `/compact':
Please organize the context of the current mandate and retain only the objectives, core constraints, confirmed decisions, modified documents, validated results, unresolved issues and next steps. Delete the overturned formula and repeat the discussion。
If the task has been split into two separate issues, it is cleaner to use a new session or `/fork ' than to continue stacking in a chat. A session would best revolve around a coherent goal, rather than a front end in the morning, an analysis of stocks in the afternoon and a continuation of the Bug at night。
XII. Modelling and reasoning intensity: Selecting by task difficulty, not using the highest level as default
Model selection is the most easily influenced by marketing. At first, I felt that it would be best if the strongest model and the highest degree of reasoning were always chosen. When actually used, it was found that small tasks would slow down, with a larger number of long answers and not necessarily a significant improvement in quality。
My division is:
- Document interpretation, fine-tuning of text, clear single document modification: low priority or medium reasoning。
- Multi-document function, conventional re-construction, test completion: medium or high reasoning。
- It is difficult, Bug, structural migration, long-link missions, security clearance: high or higher reasoning。
The model is selected on three dimensions: whether the task is complex, whether the cost of the error is high, whether you can quickly verify it. A button text modification does not need to be configured at the slowest; the migration of production data, even with few codes, deserves more careful reasoning and review。
adjustments to the current session through `moder ' and `/researching ' are more appropriate for the learning phase than frequent changes to the global configuration. when you know what type of task you're in, write the stability preference in ~/.codex/config.toml'. the global configuration saves the individual default, the project level `.codex/config.toml ' saves the specific settings of the repository, and the one-time command line parameters only deal with temporary exceptions。
XIII. MCP, Skills and Plugins: Run manual processes and automate

I USED TO INSTALL A LOT OF MCPS AND PLUGINS AT ONCE, AND THE RESULTS LIST BECAME LONG, AND REQUESTS FOR AUTHORIZATION INCREASED, AND WHEN I ACTUALLY DID THE JOB I DIDN'T KNOW WHICH TO CALL. I THEN MADE A RULE FOR MYSELF: ** THERE WAS NO MANUAL REPETITION OF THE PROCESS THREE TIMES AND THERE WAS NO AUTOMATION. **
MCP is suitable to connect to an external system, allowing Codex to access real-time information outside the chat box and the warehouse. For example, read the draft, see error monitoring, search the worksheet or access the internal knowledge base. It addresses the question of where the context is。
Skill is suitable for encapsulation repetition. It combines the trigger conditions, steps, input output and necessary scripts so that the long hints are not repeated. It addresses the question of “how this should be done in a stable manner”。
The plugins typically package Skills, MCP and associated capabilities and are suitable for the installation of a set of capabilities related to a service or workflow. It addresses the issue of “how to more easily distribute and activate a set of capabilities”。
Three questions can be asked as to whether access is worth it: whether the data is outside the project; whether the data are constantly changing; and whether the connection will save the possibility of continuous replication of paste. If all three answers are not, the ordinary documents and the hints are sufficient。
The most rational sequence of upgrades to a newcomer is to complete a purely local project; to write `AGENTS.md ' ; then to make a re-emergence review or release process into a Skill; and finally to access only one external tool that is really needed. It's a layer of power, not a pile of it。
Fourteen, ten pits I stepped on, and more direct modifications
- Start in error directoryI don't know. Amendments: Check the work catalogue and list of documents first after start-up。
- One sentence requires the completion of large projectsI don't know. Amendment: Define the minimum closed ring and break the target into a round of acceptable tasks。
- No Git checkpointi don't know. amendments: a mission is submitted once and diff is examined after completion before a decision is made to retain。
- Just describe what you want. You can't change anythingI don't know. Amendments: Clear dependencies, interfaces, directories and data security boundaries。
- "Having success" as "demand completion"I don't know. Amendments: Write the receipt and inspection list using a user action and run the test is only one of the evidence。
- Found Bug and made it large-scaleI don't know. Amendments: Require root-cause analysis before minimum modifications are made and final re-entry check。
- From the beginning, you're given full authorityI don't know. Amendments: Read-only exploration, work area implementation, high-risk actions identified separately。
- Put all the rules in every hintI don't know. Amendments: Long-term rules are written in `AGENTS.md ' , which repeats the process into Skill。
- The same long session handles unrelated tasksI don't know. Amendments: A session corresponds to a coherent goal, organized and compressed too long, and opened fork。
- Trust in the final summary, without the actual differenceI don't know. Amendments: View the modified file, command output, test results and Git diff. Agent's statement is not evidence, but the execution record。
Fifteen, five introductory phrases that can be copied directly
1. Quick Understanding of Strange Projects
Do not modify a document first. Please identify from the current project: the purpose of the project, the main directory, the start-up entrance, the reliance on management, the build/test command and the three areas most likely to have problems. Each conclusion is documented. Finally, give me a shortest path from zero to start a project; if there is not enough information, make it clear what is missing, do not guess。
2. Translating vague ideas into enforceable needs
I'd like to do it. Do not write a code, but consider the needs, like the product manager and the engineer: indicate the target user, the core scene, the minimum functional closed loop, a clear range of not to do, key risks and verifiable acceptance standards. Up to five key issues affecting the programme were raised with me; after I had responded, a phased implementation plan was exported。
3. Safely achieve a function
Target: [Function]. Context: [relevant documents/faults/references]. Constraints: [Non-changed interface, dependence, directory and style]. Completion standard: [test, build and user operations results]. Read the document and give a maximum of 6 steps plan; make only the minimum changes related to the target; run the relevant validation after completion, listing the modified document, the command results and the remaining risks。
4. Recoverable restoration Bug
Problem phenomenon: [actual results]. Retrogressive step: 1. [step] 2. [step] 3. [step]. Expected outcome: [what should happen]. Please first locate the root causes and indicate the evidence, and do not reconstruct them immediately; then propose a minimal repair programme. It is then validated in the same set of steps, supplemented by tests or inspections that prevent return. Do not modify unrelated files。
V. Self-examination before the end of the mandate
A delivery review is performed before the end: article by article against the original target and acceptance criteria; see if Git diff contains irrelevant modifications; run tests, build, Lint or type checks related to this change; check error handling and boundary scenes. The final output is only four parts: completed, validated, not completed, risk and next step. Items without evidence are not marked as completed。
Think after
After some time, the biggest change in Cordex's perception was that the hint was just the mission entrance, and the real decision was the work system。
A reliable system includes a correct catalogue of work, recoverable Git checkpoints, sufficiently clear targets and boundaries, just enough authority, acceptance and inspection standards that can be implemented, project rules that are continuously updated, and post-mission testing and review. The stronger the model, the more important the system is, because the powerful Agent can magnify the right direction more quickly and the wrong assumptions more quickly。
So the new guy doesn't have to go after "Let Codex produce the whole product at a time." Let it do a small task in the correct catalogue, produce verifiable evidence and gradually expand the mandate. You can see why it's changed, how it's withdrawn and how it's done。
If you remember only one sentence, you can remember this formula:
> Codex 's delivery quality = a clear target x a valid context x a reasonable authority x an enforceable acceptance。
ANY ONE IS CLOSE TO ZERO, AND THE END RESULT WILL BE A DISCOUNT. YOU KNOW, WHEN YOU GET THESE FOUR, YOU GET NOT JUST AN AI THAT CAN WRITE CODE, BUT A WAY OF WORKING THAT CAN CONTINUOUSLY EXPAND PERSONAL EXECUTION。