At 09:08 this morning, I sent two screenshots to a group of friends. My review had two words: “Codex. F***s.”
The screenshots did not show a new feature or benchmark. They showed one Codex task saying that it would contact another task, carry out a job, and report the result to the task that managed the project.
For months, I had done that routing myself.
The old work
Before the Codex desktop app, I worked in Claude CLI with many terminal windows. One session knew the updater. Another knew the optimizer. Others handled drivers, remote access, the app, and reviews. Each could do serious work, and its own history remained in its window.
The links between those sessions did not.
When one session found a fact that changed another session’s work, I selected the useful text, found the right terminal, and explained why it mattered. A green check could become stale because another window held a later decision. After a break, I had to rebuild the project state before I could answer a plain question: can this release ship?
For almost two weeks, I have worked almost only in the Codex app. Named tasks and projects made it easier to leave and resume work. The larger change came when I gave one task a different role.
I told it: you are the project manager.
It did not own a feature. It owned the current goal, the active work, the dependencies, the reports, and the release gates. Existing worker tasks kept their own jobs and context. They sent short reports to one place.
Then the reports began to arrive.
The morning a release stopped
The optimizer was moving into its own package. The live system already ran version 1.3.1-beta.1, but the new package called itself 0.1.0. A person can understand why a new package might restart its version number. An updater sees an older number and may refuse the install.
The optimizer task found the mismatch. The project-manager task sent the result to the updater and release owners. The optimizer chose 1.3.2. The release stayed closed until the dependent tasks agreed on the new contract.
Other reports tested the limits of our claims. A live Sungrow inverter proved one hardware path, but it did not prove the GoodWe path. Regression tests were not the same as a GoodWe device. A beta image existed, but its evidence step had failed. Since the tag was meant to stay fixed, a normal rerun could have replaced an artifact that should never change. An updater review passed, then a later review found timing faults. Even after the code became clean, a live pilot still had to pass.
The useful output was not more code. It was a set of exact stops:
* no release on the wrong version;
* no hardware claim without that hardware;
* no rerun over a fixed image tag;
* no merge before the live pilot.
That is when the workflow became real for me. A fact did not only appear in a chat. It reached the owner whose work it changed and kept later work closed.
The new term came later
People have started calling this graph engineering. The term is new and has no fixed meaning. OpenAI talks about tasks, multi-agent workflows, delegation, and orchestration, not a product named graph engineering.
The graph description still fits my work. A task owns one result. A report links it to another task. Code, tests, a live device, or a source supplies proof outside the agents’ own agreement. A gate states what must pass before work continues.
The tasks do not share one mind or one context. The work lies in making each transfer clear.
I also learned the limit. Four days before this release work, I opened a task called “Orchestrate everything.” The agents produced useful changes, but the active set grew beyond what I could inspect. I stopped new work and asked for an overview. The project task archived 39 old or overlapping tasks.
More tasks had not given me more control. A smaller active set did.
The workhorse still needs an editor
The Codex app makes this structure visible. Tasks keep their names, histories, diffs, and project place. Worktrees separate parallel code changes. One task can retain the project view while workers research, build, test, and review bounded parts.
GPT-5.6 makes the method useful in practice. It has become a real workhorse for me. It reads large repositories, uses tools, meets failed checks, accepts review, changes its plan, and continues. It can also stop when the evidence does not support the claim.
This episode tested the same process. GPT-5.6 read the podcast repository, my recent task history, the screenshot, and current sources. It wrote a script. We rendered 40 minutes of cloned voice, made six ElevenLabs music beds, mixed the audio, checked the loudness, and ran speech-to-text checks. Every mechanical check passed.
I listened and said: the intro is missing, and the episode is boring.
The first cut had the facts but not the experience. This second cut follows one morning, lets the terms arrive after the story, and gives the real intro its proper place. The model did much of the work. Human judgment decided that the work was not yet good.
A Monday version for a small team
For my cofounders, I would start small. One project task holds the outcome, decisions, owners, risks, and stop rules. Two or three worker tasks get clear contracts. One integration owner combines approved changes. Risky work gets an independent review. Active code work stays within what we can inspect, usually three to five tasks at most.
At lunch, the project task should answer four questions in a short paragraph: what are we finishing, what changed, what is blocked, and what decision belongs to us? At the end of the day, lasting decisions leave the task and enter the repository, roadmap, or another record the team owns.
Agents can build options, test them, and show where our evidence ends. They cannot decide what our company should become or which promise we should make.
The project manager can be a thread. The purpose is still ours.
Sources
* GPT-5.6
* Codex-maxxing for long-running work
* Graph Engineering guide from AI Builder Club
* From Loop Engineering to Graph Engineering?
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit frahlg.substack.com
Fler avsnitt av Coordinated with Fredrik
Visa alla avsnitt av Coordinated with FredrikCoordinated with Fredrik med Fredrik Ahlgren finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
