Slinging work instead of writing code

Cian Clarke
18 Jun 2026
  • Share

An engineering leader’s perspective on the transition from writing code to crafting an agent, to orchestrating a fleet of agents — what worked, what broke, and why it's not as scary as the discourse suggests.

If Spec-Driven Development (”SDD”) was the last real step-change in how we build, orchestrating swarms might be the next. In more than two decades of building software, the last time the ground moved like this for me was going from PHP to Node.js v0.6. On a recent week off, I got to spend some rare time actually building something — and applying many of the techniques I’ll discuss below to build something real and of substance.

The idea was a tangent — could I develop a low-latency arbitrage trading bot for prediction markets, combining a language I knew nothing about (Rust) with a domain I knew nothing of (sports betting), using Claude Code and agent teams to do the heavy lifting?

The driver was simple — there's no doubt techniques like SDD allow us to move much faster as engineers, but was there another level of scale available to us? That is the promise of orchestrating swarms of agents, and what follows is what worked, what broke, and why it's not nearly as daunting as the discourse suggests.

Plan mode hits a wall

I went in with a hypothesis: Claude Code's harness has become so good that I could likely combine plan mode with tactful prompting to make my way towards a working monolith deployed on EC2 without writing extensive specs along the way.

In practice, I was orchestrating anywhere from 2 to 5 paned Claude Code windows in iTerm2, and sometimes one of these panes would spawn its own smaller swarm of workers using Claude Code’s native teams functionality. There was no orchestration layer between tabs, and no shared understanding of what had already been done.

This didn’t pan out as I’d hoped, and many of the early vibe coding problem areas arose again:

  • Lots of repeat regressions — things breaking that had previously worked fine. Context and memory getting lost between sessions.
  • A work log per task would have helped, but I was eager to see how far the bare Claude Code harness could be pushed.

Having a better way to hand off tasks to the agents seemed like a natural next step, and Beads fit the bill.

Beads: units of work for agents

Beads are work items (akin to stories or tickets), but designed specifically for agent workflows. Think of a system like Jira or Linear, tailored for agents, with structured storage of work items, explicit dependency management, and a persistent change history.

Beads did help steer the model across tasks. Having a persistent record of what's been done, what's in flight, and what depends on what made a noticeable difference to the regression problem.

There were still some rough edges, though:

  • Filing new beads for every task was a pain. Each one needs a title, description, type, priority — you can have Claude guess, but it adds friction.
  • Sometimes Claude would start work automatically on a bead. Sometimes you had to explicitly tell it to begin by referencing the ID. This felt a bit like Russian roulette.
  • Claude would often just skip filing the bead if you weren't really explicit about it and start working anyway. (Even at the frontier, agents not following instructions is a problem. Some things never change.)
  • Work threads would sometimes get stuck waiting for a decision from me, the pesky human, and get hung up for a long period.
  • The main Claude work loop was tied up and busy when you slung work. You had a single async thread to ask questions in, but you couldn't move on to instructing new work — it'd queue the message.

I felt I was starting to notice behaviours that Gas Town was designed to solve, so it was time to take that puppy for a spin.

Gas Town: where things clicked

Switching to Gas Town changed the experience significantly. For the uninitiated, Gas Town is Steve Yegge's multi-agent orchestration rig, introduced in the irreverent, over-the-top article “Welcome to Gas Town.” You interact with a "Mayor" who coordinates work across a team of workers called "Polecats." Some of the constructs in Gas Town cascade down into Claude Code, which then cascades further downstream into sub-agents. It's turtles all the way down.

Filing beads with the Mayor is easy — the bead substrate fades into the background, which feels far more intuitive. The issue-tracking system (designed for agents, not humans after all!) fades into the background, letting you focus on best describing tasks.

The shift is that you're no longer telling an agent to do something; you're slinging work into a system, and the system figures out how to schedule it.

On rare occasions, the Mayor would still decide "it'll be quicker for me to just do this myself," and you've got to Cmd-C that immediately, because it's a complete anti-pattern. Using the word "sling" in every instruction reliably avoids this in my experience — funny how a single verb can change an LLM's behaviour so much.

The Handoff pattern solves the work loop problem I had with raw Claude Code. When a Polecat gets blocked or finishes, work flows naturally rather than getting stuck.

When a task is stuck awaiting human input, a nudge will often get the Polecat back on track.

Add the intra-agent messaging substrate — where workers can coordinate without stepping on each other's toes — and you wind up with something that feels genuinely like orchestration rather than babysitting.

I was typically running 2 to 7 tasks in parallel. I haven't found the need for higher degrees of concurrency on a single repo, and I'm not yet sure how people get to 10+.

Self-healing software (sort of)

Gas Town has a "Formula" + "Molecule" concept where you define reusable patterns for common types of work. I started using this construct to define a reusable task that could keep an eye on logs and run on a cron. The system would then spot opportunities to autonomously fix issues — a sort of self-healing software.

I had the bot running, a Formula watching the logs, and when something went sideways, the system would detect it, file the work, attempt a fix, then sling a deployment. This got pretty close to the point where I’d have the confidence to put it on full autopilot.

This was also the first time I've seen something that feels like genuinely self-healing software.

The rough edges (and Gas Man)

There are still lots of Gas Town bugs. The system knows so much about its own internals that it often fixes itself and relinks the local binary — a good advertisement for the self-healing concept, but it’s also not strictly producing fixes suitable for merging upstream, so I’m not sure how I feel about this.

The biggest gap I found was visibility. The launch blog discourages you from following along with Polecat logs, but there is no great dashboard view of work: neither gt feed nor gt dashboard shows live Polecat logs performing work.

I still want to be able to glance at the live work log of ongoing tasks. I was frequently spotting agent rabbit holes that I knew had a faster solution.

I wound up building my own tmux-based Mayor dashboard: Gas Man. ("Gas Man", as in, ah sure he's a gas man...)

When interacting with the Mayor, which spawns new work for Polecats, a new tmux pane docks to the right of the Mayor window. You get a live view of everything that's happening.

There are still loads of situations where human intuition wins over the agent. One that sticks out: SSH was hanging when connecting to the EC2 instance. The agent started going down a rabbit hole — maybe the instance needs a reboot, maybe there's a networking issue, let me check the logs... nope, it's clearly a misconfigured Security Group. Took me about three seconds to spot, replacing what would have taken the agent five minutes of flailing.

It still needs some polish and would probably be better directed towards the Gas City ecosystem (the follow-up to Gas Town) at this point. Gas City hit v1.0.0 in late April, a milestone some have been waiting for before jumping ship. The time, as they say, is nigh.

The safety question

I'll say it plainly: I did not see any of the scary failure modes people talk about. No repo deletions. No runaway infrastructure. No p5.48xlarge instances appearing on my AWS bill. Nothing particularly harmful. The system seems more bounded than the discourse would have you believe.

I certainly hit usage caps — in the beginning, this took longer than expected — but as parallelism increased, I was quickly burning through one Claude subscription and two separate Codex subscriptions in a week.

Purely anecdotally, I pushed far more code through Gas Town with Claude Code than with a similar seat type in BMad (a popular spec-driven development framework) with Cursor before hitting usage caps. Multiplexing subscriptions from Claude Code to Codex worked well — outputs are indistinguishable, but the Codex-Gas Town integration has far more rough edges.

Where this leaves us

My attempts have not resulted in new application-driven riches, but along the way, it does feel like I've stumbled upon a level-step change in how software is made. Like I said up top, this feels like a bigger shift than any I've seen before.

The idea of overseeing a factory of workers — sometimes with the lights on, sometimes in the dark — feels like a truly novel approach to the craft. You move from writing code step by step to orchestrating parallel units of work. You spend more time reviewing, directing, and making architectural calls than you do typing syntax. And at this scale, reviewing every line of code is just no longer feasible — you need really strong CI/CD gates and review agents picking up the slack.

We hit the orchestration bottleneck at Nearform early on. The challenge of orchestrating a swarm of agents and a small team of 3-5 people looks much the same: merge queues, reliable accentuated review pipelines and a solid approach to harness engineering are the solution. It's a fair bit of what we spend our time on at Nearform. If this is a challenge you're facing, get in touch.

The projects I’ve been pushing are small enough — hovering around 15k LOC — mostly Rust, some IaC, some Python, and some Bash. I wish I'd started with something more ambitious, but even at this scale, the difference in operating model was clear.

A lot of the prior sentiment around the four rules of Gas Town, all being variants of "do not use Gas Town," is largely gone. The primitives can be off-putting at first — there's a lot of new vocabulary to absorb, and Mad Max / steampunk naming conventions are not to everybody's liking — but the Mayor guides you through most of the paradigms, and you can very much learn as you go.

I think a lot of people are going to be surprised by how quickly this way of working goes from "interesting experiment" to “how we actually build things.”

But wait - there's more.

Nearform publishes real-world learnings on data & AI, engineering, and digital strategy - with more merged in weekly.

Insights

Perspectives on AI in engineering, product development, and strategy, for enterprise executives.

Community

Deep dives and tutorials by engineers, for engineers.

Insight, imagination and expertly engineered solutions to accelerate and sustain progress.