Contents
Picture a small building site with no foreman: every worker hauls bricks wherever they like, the walls do not meet, and “almost done” is the hourly chorus. That is what a day with several AI agents often looks like without a human who sets bounds and accepts the work. With a foreman the picture changes: one engineer holds the plan, assigns tasks, checks the joints, and decides what may ship.
Not long ago a complex product almost automatically meant a large team: architect, frontend and backend, testers, DevOps, security. Today part of that work can be delegated to AI agents — and one experienced programmer can organize them the way they once organized several specialists. This is not magic that “the model will assemble anything.” What changed is different: the developer now has digital workers with limited rights, who can be given narrow tasks whose results fit into a system.
The main advantage is no longer typing code faster. It is stating tasks well, designing the system, coordinating agents, and telling a finished solution from a draft that only looks convincing.
Key takeaways
One strong engineer with agents widens their radius of action; they do not outsource responsibility. Models write code, run checks, and prepare drafts. A human sets bounds, accepts risk, and signs the release.
An agent crew is roles and contracts, not five different products. Developer, tester, reviewer, documenter, and researcher can be separate runs of the same tool with different tasks and access to context.
Experience beats typing speed. The more complex the system, the more the person who sees architecture, can decompose work, and already knows what to verify wins.
Throughput beats the count of parallel agents. Coupled tasks, file conflicts, and weak acceptance eat the gain from parallelism. The goal is useful process speed, not a record number of tabs.
Without rights, tests, and done criteria, an agent is a risk, not an accelerator. Security, migrations, payments, and personal data need procedures — not “trust the model.”
From chat with prompts to a crew of workers
The familiar way of working with AI is simple: ask a question, get a code fragment, fix mistakes, continue. That helps, but it still keeps the whole process in one head and forces constant context switching.
An AI agent works in a wider loop. Depending on the tool, it can explore the repository, change files, run tests, use the terminal, look at the result, and take the next step without a separate command for every action. It still does not become a fully autonomous employee: its reach is bounded by tools, context, permissions, and how well the task is framed.
The next step is not one universal assistant, but several workers in parallel. On a building site those are different trades; in a repository — different roles:
- Developer agent implements a single feature or module.
- Tester agent drafts scenarios, checks edge cases, and runs tests.
- Reviewer agent looks for bugs, duplication, maintainability issues, and potential vulnerabilities.
- Documenter agent updates the README, API description, and developer guides.
- Researcher agent studies existing code, compares implementation options, or prepares a technical plan.
That split does not have to mean five different models. Often the same tool is launched several times with different tasks and access to context. What matters is not headcount, but a thoughtful organization of work — as on a site the schedule and the joints matter more than how many people hold a shovel.
I covered the related frame “agent = model + process scaffolding” in pieces on agentic engineering in 2026 and the new development cycle: vibe coding and agentic engineering.
Why an experienced developer gets more leverage
At first glance it seems AI helps beginners most: describe an idea in plain words — and here is a prototype. For small sketches that can work. But the more complex the system, the more knowledge that is not reducible to code generation matters. A foreman without a drawing does not speed up the build — they speed up chaos.
They see the system as a whole
A function can work on its own and still break the architecture. An API change ripples into client modules. A new table violates constraints. Speeding up one query worsens behavior under load. The agent will propose an implementation; the human judges it against data, contracts, compatibility, and future change. Where a model fails on mature code — in a breakdown of a 15-year legacy.
They can split a large task into independent parts
“Build an enterprise ERP” is too large and ambiguous. Framed that way, the agent produces an impressive prototype that is hard to grow. An experienced developer turns the goal into checkable stages: API contract, migration, service, interface, tests, access rights, integration. Some steps can run in parallel; some wait on earlier results. That is crew management: parallelize independent work, do not launch as many processes as possible.
They know what must be checked
Code that looks plausible is not necessarily correct. An agent can forget a null, misread a business rule, skip an access check, or write a test that confirms the same bug as the implementation. So done criteria are set in advance: which tests are mandatory, which interfaces must not change, which security constraints are critical, what happens on failure. Checking after an agent patch is its own discipline: code review in the AI era.
They steer the work, not only accept the result
When an agent errs, an experienced engineer localizes the cause faster: too little context, wrong level of abstraction, conflicting requirements, or a task that is too large. Instead of endlessly repeating the prompt, they change conditions, tighten the contract, or split the work. AI does not erase the value of experience — on many tasks it widens the reach of a specialist who already knows how to make technical decisions. On the wider shift of the profession — in “The programmer in the age of AI”.
How this looks on a real project
You need to add a request-management module to an existing system: shared backend, shared auth, database, user interface. Work without a foreman turns into five incompatible “almost done” piles. With a foreman the cycle looks like this.
Step 1. The human sets the bounds. Studies the architecture, records requirements, constraints, business rules, and done criteria. Writes a short technical plan: entities and API, files that may change, components that must stay compatible.
Step 2. The researcher agent studies the repository. Finds similar modules, conventions, integration points, and affected areas. The developer checks the findings rather than taking them on faith. Without indexing and context the model “fills in” the system from generic patterns — the same theme as in pieces on semantic context and LSP.
Step 3. Tasks are assigned to workers. One agent prepares the backend and API, another the interface against an agreed contract, a third the tests. If interface and API cannot be built independently, fix the contract first or run a short preparatory stage.
Step 4. Results go through independent checks. An agent may review another agent’s code, list risks, or propose extra tests. The final call on correctness stays with the developer. Critical checks are the project’s tests, linters, static analysis, and CI.
Step 5. The human integrates the work. Resolves conflicts, checks consistency, runs the system, walks real scenarios, and decides whether the result is ready to ship.
In this scheme the developer does not perform every operation by hand. They run the process and spend more time on architecture, task framing, risk, and acceptance. The platform around agents (tracker, environments, data, tools) often matters more than “one more fast model” — see agents need a platform, not just speed.
Economics of development: not agent count, but throughput
The “one engineer — a crew of workers” model has several advantages.
Fewer switches between kinds of work. An agent explores code or prepares a first cut of tests while the developer thinks about architecture. Not every wait turns into idle time.
Parallel work on independent tasks. Several agents study different parts of the project or prepare isolated changes at once — when ownership boundaries are clear.
Faster feedback. Early drafts, tests, and reviews help catch a problem before the solution grows unnecessary code.
Cheaper experiments. It is easier to compare two approaches, prepare a prototype, or explore an unfamiliar library — without turning every experiment into a separate large project.
But development does not speed up in proportion to the number of agents. Coordination, checking, and integration also cost time. If tasks are tightly coupled, workers interfere, duplicate work, or produce incompatible solutions. Studies of human versus model speed often disagree precisely because they measure different things: generating a fragment versus shipping into a mature repository — see human vs AI: who writes code faster.
The right goal is not the maximum number of agents, but maximum useful throughput of the process: how many checkable changes reach users without growing incidents and debt.
What you cannot hand over with unconditional trust
AI agents err, misread requirements, invent nonexistent APIs, miss rare scenarios, and confidently declare a task done when checks have not passed. Be especially careful with security, data migrations, payments, personal data, and critical production systems.
A reliable process rests on a few rules:
- Least necessary rights. The agent gets access only to the files, tools, and environments the task needs. On isolation and risk — in Docker sandboxes for AI agents.
- Checkable completion criteria. “Code is written” is not “the feature works.” You need tests, defined scenarios, and an observable result.
- Independent verification. Agent review is useful, but it does not replace tests, static analysis, human review, and CI gates.
- Clear change boundaries. Small tasks and isolated branches reduce conflicts and make rollback easier.
- Control of important decisions. Architecture, secrets, data, and production release need approval procedures — not “the model said everything is fine.”
Practice with long-running agent processes shows you need an organized environment: clear instructions, a checkable task list, a way to run the app and tests, a way to preserve progress between sessions. Otherwise even a strong model loses context or prematurely treats the work as finished. External landmarks on agent “harnesses”: Anthropic materials on harnesses for long-running agents and OpenAI documentation on orchestrating multiple workers.
A new skill — engineering management of AI
The developer gains an extra layer of work. They design not only the software system, but also the process in which AI performs part of the engineering actions. This is no longer “writing good prompts,” but engineering the interaction of humans, models, tools, and infrastructure.
The skill includes:
- stating tasks through outcome, constraints, and acceptance criteria;
- managing context: needed files, documentation, project conventions, and technical decisions;
- decomposing work and mapping dependencies;
- choosing the level of autonomy for the concrete task;
- organizing testing and independent verification;
- managing rights, cost, time, and risk;
- recording decisions and progress so work can continue safely.
On a building site that is keeping a work journal and not letting the crew into a load-bearing wall without a drawing. In a repository — the same thing in the language of git, CI, and contracts. When context is gathered poorly, the model errs with confidence; when it is gathered well, it speeds routine without breaking architecture.
Will one programmer replace a whole team?
Sometimes one specialist does ship a product that would once have needed several developers — when the project is bounded, the architecture is clear, and the work cuts into tasks with clear outcomes.
But the formula “one person replaces a whole department” oversimplifies reality. A large system includes business requirements, support, operations, security, integrations, talking to users, and ownership of consequences. AI agents take on part of the operations; they do not remove the need to decide and answer for the result.
It is better to say it another way: an experienced programmer gains the ability to direct a larger volume of engineering work than they could do alone. The bottleneck shifts from typing speed to the quality of task framing, process organization, and result checking. The advantage belongs to whoever turns an unclear idea into a clear requirement system faster, delegates safely, and ships checkable changes.
Typical mistakes when working with multiple agents
A first task that is too large. “Build the whole module” almost always yields a pretty but brittle draft. Start with a contract and one vertical slice.
Shared files without coordination. Two agents editing the same service produce conflicts and incompatible styles. Ownership boundaries first, then parallelism.
Agent review instead of tests. A second agent catches style and some bugs, but does not replace running scenarios and a human eye on business rules.
Rights “like a developer.” Full access to secrets and production for convenience is a direct path to an incident. Least privilege is cheaper than any minute saved.
The metric “how many agents are running.” That is bustle, not outcome. Count accepted changes, time to green CI, and number of rollbacks.
Where to start this week
You do not need to build a complex multi-agent platform at once. Take one real slice of the project and a short cycle — like a training day on site with one wall, not a whole block.
- Ask the agent to study a module and draft a plan without changing files.
- Pick one small task with clear done criteria.
- Assign implementation and ask it to show changed files and the rationale for decisions.
- Run the tests and review the changes yourself.
- Only then add a second agent — for testing or review — and compare whether the process got better.
That is how working practices accumulate: from help on separate operations to managing parallel tasks and then to longer agent processes. If you want to train typing speed and accuracy on real code from the site — open the code typing trainer.
Frequently asked questions
How does an AI agent differ from ordinary chat with a model?
Chat answers a question and returns text. An agent works in a loop: looks at the repository, changes files, runs commands, reads the result, and continues. The difference is access to tools and the length of an autonomous step — not “a different kind of model magic.”
Do you need five different models for five roles?
No. Often one tool with different tasks, different context, and different rights is enough. Role is set by framing and bounds, not by the logo on the tab.
What is a safer start for someone new to agents?
Tasks without access to secrets and production: exploring a module, a draft of tests, documentation, a small refactor in an isolated branch. Done criteria and tests — before any parallel work.
When do parallel agents hurt more than they help?
When tasks share the same files, the contract is not fixed yet, or acceptance boils down to “looks fine.” Independent pieces and explicit criteria first, then parallelism.
Will this replace junior developers?
AI automates part of the tasks people used to learn on. Value grows for those who can set tasks, check results, and hold architecture. Entry into the profession gets harder, not disappears — more in the article on the programmer in the age of AI.
What if the agent confidently lies about an API?
Do not argue in chat forever. Narrow the task, give precise context from the repository, demand a pointer to a file or test, and verify the claim with a project tool. A confident tone from the model is not proof.
How does this connect to enterprise processes?
Through the same gates as for people: branches, review, CI, limited rights, a decision log. An agent is another worker in the existing pipeline, not a bypass of procedures.
Further reading
Inside stuzhuklab, related pieces on agents, context, and quality help:
- Agentic engineering in 2026
- AI agents need a platform, not just speed
- The new SDLC: vibe coding and agentic engineering
- Code review in the AI era
- The programmer in the age of AI
- Human vs AI: who writes code faster
- Can AI understand a project with 15 years of history
External landmarks on orchestration and long-running agents:
Conclusion
AI agents do not make engineering thinking unnecessary. The more actions you can delegate, the more it matters to understand which system we are building, why it must work that way, and how to prove the result is correct.
The experienced developer of the future is not necessarily the person who writes all the code by hand. They are an engineer-foreman: sets direction, distributes work among tools and agents, keeps the architecture whole, and ensures the quality of the final product. On this site we treat AI-assisted development as a practical engineering discipline — from how models work and how context is organized to testing, architecture, and working processes. That approach lets you use what AI offers without giving up control of the code and the outcome.
This week one module, one agent with a plan and no edits, and one task with strict acceptance are enough. When the wall stands straight — you can call in a second crew.
Who runs this lab and which projects sit behind the practice — in the resume.



Comments