№ 02
6 min read

The first time ChatGPT saw my code

I said the next entry would be about Codex.

It is.

But when I tried to find the beginning, I discovered that I had remembered it wrong.

At the end of 2024, after a period of freelancing, I joined a new company as a frontend developer. The job title said frontend. The work did not always stop there.

ChatGPT moved with me.

By then, it stayed beside my editor while I worked. I used it to understand unfamiliar code, compare approaches, think through architecture, and debug whatever had stopped working that day.

The process was not sophisticated.

I would copy an error, paste the function around it, and ask:

What is the error? Fix it.

That was almost the entire prompt.

ChatGPT would explain the syntax, suggest a change, and give me a new block of code. I would paste it into the editor, run the application, and see what broke next. If I did not understand the answer, I would return with a smaller block and question it line by line.

It worked more often than that process deserved.

For a contained problem, GPT-4o was remarkably useful. It could spot a missing condition, explain an unfamiliar pattern, or show me a cleaner way to write a function. Even when the first answer was wrong, it often helped me reach the right question faster.

But a function is not an application.

The problems I was working on rarely respected the boundaries of one code block. A page depended on a component. The component called a hook. The hook used an API client. The response came from the backend. Three or four files could be individually correct and still fail when they met.

To explain one bug, I had to paste the first file, describe where it connected, paste the second, remind ChatGPT about the first, and then squeeze in the part of the third file that seemed relevant.

Large files made it worse. By the time I had supplied enough code, the useful part of the conversation was buried under the explanation required to begin it.

I was saving time on the answer and spending it again on the context.

Experience helped. I became better at guessing which function mattered and which file ChatGPT needed to see. We usually got the work done. But every conversation began with the same manual transfer: select, copy, paste, explain.

Then I installed the ChatGPT app on my Mac.

I expected the website in a smaller window.

Instead, I found a feature that could work with the application open beside it. I connected it to my editor, opened a page.js file, and asked a question that had nothing to do with a bug:

Tell me what you understand from this page.

I wanted to see how much it could actually read.

It described the purpose of the page, the main flow through it, and how the pieces fitted together. I had not copied a function. I had not pasted a single line into the chat. I had not explained what the page was supposed to do.

The code was still sitting in my editor, where it belonged.

I read the answer again, opened another file, and tried to catch it out.

Then another.

I opened the related frontend files in VS Code and the backend files in PyCharm. I arranged the panes so the relevant code was visible and gave ChatGPT access to them. Now I could ask about the connection between files without rebuilding that connection inside the prompt.

The feature did not understand my entire codebase. It could only read the editor content I had opened and exposed to it. But that was enough to change the rhythm of the work.

After months of carrying fragments of the application into ChatGPT, I could leave the code where it was and bring the question to it.

Debugging became faster. Understanding an unfamiliar page became easier. When I moved between the frontend and the backend, I no longer had to flatten both halves into one enormous message before I could ask how they fitted together.

That particular improvement had not come from a new model.

I had stopped starving the model I already had of context.

Context, however, was only one thing I was learning to manage.

My Plus subscription gave me a model picker that kept changing. At first, I used GPT-4o for almost everything. Then o1 arrived.

o1 deserves its credit. It was one of the best models I had used. I did not waste its limited messages on every syntax error. I saved them for architecture, trade-offs, and decisions where the first plausible answer was not enough. It was slower than GPT-4o, but for those questions, waiting was part of the value.

For everyday coding, GPT-4o remained the fast loop. When the code needed more thought, I used the higher-reasoning options in ChatGPT: first o3-mini-high and later o4-mini-high. I did not spend much time with the standard o4-mini option. The high option was the one that earned a place beside GPT-4o in my coding workflow.

Then came o3. For me, it felt like an upgrade: the next step after o1, stronger on the architectural and multi-step decisions for which I had been saving my limited messages. o1 had taught me to wait for a model to think. o3 made that habit more useful.

By then, the model picker no longer looked like a leaderboard. It looked like a toolbox.

I was no longer asking which model was best.

I was asking what kind of thinking the next question needed.

For a while, this felt like the missing piece. ChatGPT could see enough of the work to discuss the system rather than one isolated function. I still reviewed the answer. I still made every change. I still ran the application and decided whether the result was correct.

But the clipboard was no longer standing between the conversation and the code.

Months later, when I learned about Codex, I looked back at this experience and joined the dots too quickly. The Mac integration had felt like an early version of Codex, so I began thinking that I had used Codex before it was officially released.

That is a neat version of the story.

It is also wrong.

The feature was called Work with Apps. It passed content from my open editor panes into ChatGPT. It was not an agent working across my repository. It was not running commands, changing several files, executing tests, or checking its own work.

It could see.

I was still the hands.

The name matters because the difference matters. Work with Apps removed the need to carry every useful fragment into the conversation. Codex would go much further: it would enter the repository, make the change, run the code, and return with evidence of what happened.

The first boundary was context.

The next one was action.

That is where Codex begins.


— H.