The strangest thing Codex changed was not my code.
It was the application I opened to write it.
For years, opening a project meant opening VS Code. Then Terminal. Then the browser.
One day, I noticed I was opening none of them first.
I was opening Codex.
This is the second half of the story I began in The first time ChatGPT saw my code.
That story ended with a boundary.
ChatGPT could see the work.
I was still the hands.
Codex had started doing it.
The change began inside VS Code.
I cannot tell you the exact model I used on the first day. It was somewhere around GPT-5.1-Codex or GPT-5.2-Codex. The model picker was changing faster than I was keeping a diary.
What I remember is the difference.
GPT-5-Codex had been useful. GPT-5.1-Codex felt like a serious jump. GPT-5.2-Codex improved the loop again. But the important change was not another cleaner explanation or a slightly better block of code.
Codex was inside the editor with the repository.
It could search for the relevant file, follow an import, make a change, run a command, read the failure, and try again. I no longer had to carry the code to the conversation one function at a time.
The conversation had entered the work.
At first, I still treated it like ChatGPT with better access.
I asked a question. It answered. I inspected the change. Then I switched to Terminal on my Mac to run something, returned to VS Code, opened another file, and continued the conversation.
It worked, but the workflow had seams everywhere.
VS Code held the files. The terminal held the commands. The browser held the result. The chat held the reasoning. I was still moving between them and keeping the complete state of the task in my head.
Then I met AGENTS.md.
It looked almost too small to matter: a Markdown file with instructions for an agent. But it solved a problem I had been recreating in every conversation.
I had spent years onboarding developers into repositories. Now I was onboarding the agent.
This is the project structure. These are the safe commands. Do not touch this area. Preserve existing changes. Run these checks before you say the work is finished.
The prompt stopped carrying every permanent rule because the repository could carry its own rules.
I began to understand the difference between giving an AI more context and giving it the right context. The whole repository was available, but that did not mean every file belonged in the task. A useful instruction named the goal, the relevant area, the constraints, and what proof I expected at the end.
That was the first real change in how I worked with Codex.
I stopped asking only for code.
I started asking for finished work.
When it began to feel agentic
GPT-5.3-Codex was the breakthrough for me.
Earlier models could complete impressive coding tasks, but I often felt the conversation waiting for me at every boundary. With GPT-5.3-Codex, the model felt more willing to stay with the problem: inspect before editing, follow the failure, use the available tools, and return with a result rather than a suggestion.
That is what agentic came to mean in my daily work.
Not that the model wrote more lines.
It crossed more of the distance between the request and the evidence.
Around the same period, Codex-Spark gave me another kind of improvement: speed. I used it heavily for focused iterations where waiting would break the flow. Inspect the component. Make the narrow change. Check it. Move to the next problem.
Spark did not replace the deeper model for every decision. It gave the toolbox a fast tool that I could reach for repeatedly.
The model picker still mattered, but less than before. A strong model inside a weak workflow could still produce an elegant mistake. A faster model without a clear boundary could reach the wrong destination sooner.
I was learning that the agent was only one part of the system.
Then came the Codex app
Once before, I had installed a Mac app expecting the website in a smaller window. That mistake introduced me to Work with Apps.
I should have learned not to underestimate a new window.
I made the same mistake again.
The Codex app did not feel like one more place to chat about code. The tasks, repository, diffs, tools, and results lived together. I could start with a question, let Codex inspect the project, review what it changed, and continue without rebuilding the task in another application.
When the integrated terminal arrived, another seam disappeared.
I did not stop using the shell. The commands still mattered. Their output still mattered. I stopped opening a separate Terminal window for most of the work because the shell was now available inside the same environment as the conversation.
The same thing happened to VS Code.
I did not make a decision one morning to abandon my editor. I simply noticed that I was opening it less. Then rarely. Eventually, almost all of the work I had been distributing across VS Code, Terminal, ChatGPT, and a browser was being managed from Codex.
The editor had not lost a competition.
The boundaries around it had dissolved.
I looked back at my own habits
There is a dangerous stage with any new tool: the moment when using it a lot begins to feel the same as using it well.
I wanted to know whether my workflow had actually improved or whether I had only become faster at opening Codex tasks. So I looked back across the work.
I expected a collection of prompts. What emerged was a change in the shape of the work.
My earlier usage was direct and tool-light. Ask for the change. Read the answer. Run a few things myself. Return when something failed.
My current loop looked different:
Context → scope → inspect → plan → edit → test → preview → review → handoff.
That sequence sounds formal when written on one line. In practice, it is how I keep an agent from being impressively wrong.
Take a common frontend problem: a number on the page does not match the value returned by the API.
In my old workflow, I would copy the component into ChatGPT, then the hook, then part of the response, then the function that transformed it. By the time the conversation understood the bug, I had already performed most of the investigation manually.
Now I can point Codex at the route, name the behavior, attach the screenshot, and set the boundaries. Inspect first. Do not edit yet. Keep the API unchanged. Preserve unrelated work. Tell me why the values differ and what evidence will prove the fix.
Codex reads the repository instructions, traces the value through the API client and component, and proposes a plan. I correct the plan if its boundary is wrong. Then it implements the change, runs the relevant checks, opens the application in the built-in browser, and compares the rendered result with the real response. It checks the console, desktop and mobile layouts, and the Git diff. If I said no commit or push, it stops there.
The code change may be small.
The completed loop is not.
Context before code
I now begin many tasks by asking Codex to understand and discuss before it edits.
The request includes the route, API, screenshot, file, or behavior that matters. It also includes the limits: read-only for now, do not change the API, preserve unrelated work, do not commit, do not push, or use only safe request methods.
Those limits are not ceremony. An agent that can act needs to know where action must stop.
The improvement was not writing enormous prompts. It was learning which facts change the decision.
Proof after code
I no longer accept “implemented” as proof that something works.
For backend work, Codex can inspect a real response, compare keys and counts, run framework checks, execute targeted tests, and then run the wider suite when the risk justifies it.
For frontend work, the browser became part of the engineering loop. Codex can open the actual page, navigate the flow, compare the rendered values with the API, check the console, and look at desktop and mobile layouts. A screenshot can show whether the result merely exists or actually belongs on the page.
Then it can inspect the Git diff and tell me exactly what changed.
This is the part of Codex that changed my work more than code generation.
Writing code was never the entire job. The expensive part was often proving that the code fitted the system around it.
Now the same agent that makes the change can gather much of that evidence.
I still judge the evidence.
Skills instead of repeated explanations
Some instructions kept returning: how to verify an application, how to inspect a UI, how to hand off a long task, how to work with a document, how to check a change without publishing it.
I began using and creating skills for those repeated workflows.
A skill is not magic hidden behind a button. It is a procedure I no longer want to reconstruct from memory every time. It can contain the instructions, references, scripts, and checks needed to perform one kind of work properly.
The value is consistency. The second run should not depend on whether I remember the best sentence from the first prompt.
Stable project rules belong in AGENTS.md. Repeated procedures belong in
skills. The current task belongs in the prompt.
Separating those three reduced a surprising amount of noise.
More than one agent, without creating a crowd
Larger work introduced another change. One agent did not need to perform every independent investigation in sequence.
I could separate research, implementation, API validation, browser QA, and review into bounded workstreams. Separate worktrees kept parallel code changes from colliding. The main task could coordinate the result and remain responsible for integration.
This was powerful, but only when the work was genuinely separable. Five agents editing the same decision are not a team. They are a race condition with good grammar.
Memory that survives the conversation
Chat history is a poor home for architecture.
For decisions that need to survive, I use Obsidian as a durable explanation layer. Codex reads the repository, compares the implementation with the existing notes, and updates the architecture or onboarding material after the direction is agreed.
Git remains the source of truth for the code. The notes explain how the pieces fit together and why a decision exists.
This completed another loop. Codex was no longer helping only with the current change. It was helping leave the project easier to understand for the next change.
Did the hallucinations stop?
I almost wrote that the newer models stopped hallucinating.
That would be a satisfying sentence.
It would also be false.
With GPT-5.6 Sol, the experience is closer than it has ever been to a fully agentic workflow. It feels less eager to guess and more willing to inspect. It can reason across a large task, use the shell, work with files, use skills, inspect the browser, and keep moving through several stages without asking me to manually transport every result.
In my work, hallucination is no longer the constant interruption it once was. But some of that improvement comes from the model, and some comes from the system around it.
The model can inspect the real repository instead of guessing its structure. It can run the command instead of imagining the output. It can open the page instead of assuming the component looks right. It can compare the diff instead of relying on a memory of what it changed.
The answer has more ways to meet reality before it reaches me.
That is not the same as being incapable of error.
An agent can still misunderstand the goal, choose a locally sensible solution that is wrong for the product, or verify the wrong success condition. Greater agency makes review more important, not less, because a mistake can now travel through more files and more tools before it stops.
The dangerous Codex task is not always the one that fails.
It is the one that completes the wrong job perfectly.
What happened to my job
I write less code by hand now.
That sentence can sound like I am doing less engineering. My experience has been the opposite.
I spend more time defining the boundary, choosing the architecture, separating independent work, checking evidence, reviewing trade-offs, and deciding whether the result belongs in the product.
The keystrokes moved.
The responsibility did not.
Codex began as an extension beside my code. Then it became the place where the code, terminal, browser, project rules, reusable skills, and parallel tasks met.
Somewhere during that transition, it stopped feeling like a tool I opened to help with development.
It became my development environment.
In the previous entry, ChatGPT could see.
I was still the hands.
The hands can belong to the agent now.
The direction is still mine.
The editor disappeared.
The engineer did not.
— H.