A governed coding agent

AI writes the code. CodeLead decides what is allowed to land.

CodeLead runs on your machine with the model you choose, local first. It scopes every change, refuses what it cannot show is safe, validates before moving on, and leaves evidence a third party can follow. New applications, changes to an existing code base, or the modernization of a legacy system: the same discipline, whatever the task.

codelead · local 4B model · governed run
$ codelead /implement "show the page title above the game"
→ model proposes: rewrite vite.config.ts, tsconfig.json, package.json …
✗ refused by patch gate — approved change surface is index.html only.
3 files outside scope. Nothing applied. Reason recorded to evidence.
(attempt 17 of 24 — all 24 refused; the working game was never touched)
→ retry with hard constraint: scope = index.html
✓ applied · build green · evidence record written

A real run. Near the end of building a Pong game, a 4-billion-parameter local model tried twenty-four times to rewrite the project's config files to add a page title. Every attempt was refused; the game kept working. The smaller the model, the more the governance matters.

What it is

The model is replaceable. The method is the product.

Local-first, any model

Runs open models on hardware you control through LM Studio or Ollama, so your source code stays inside your boundary. Remote models are available when you choose them, and that choice is always explicit: what goes to a remote model leaves your machine.

Nothing reaches disk unchecked

The model proposes; it never edits files directly. Every change passes hard safety checks first, and a change outside the approved scope is refused with a written reason. Zero unsafe applies across 350+ benchmarked runs.

Small increments, real commits

Work is broken into increments you approve. Each one builds, passes its tests and behavior checks, and lands as a commit a person can read, or it does not land at all.

Evidence for every change

Each change closes into a record that runs from request to approval, and what it teaches lands in a knowledge base where a verified fact and a guess are visibly different things.

The flagship use case · Modernization

It is rarely because the code is old.

Teams call us because the risk has become impossible to ignore.

  • The only developer who understands the system is retiring.
  • A customer, auditor, or acquirer asked about technical risk.
  • A cloud migration is blocked by one old application.
  • The system runs on an unsupported runtime or a fragile deployment path.
  • A previous modernization attempt stalled.
  • A small change could break billing, inventory, claims, dispatch, or reporting.

The real risk

The problem is not old code. The problem is unknown behavior.

Most modernization efforts do not fail for lack of ambition. They fail because they begin with a rewrite before anyone has a reliable map.

Business rules turn up halfway through the project. A retired developer's shortcut turns out to be a critical workflow. A small database change breaks a downstream report. The new system passes the demo and fails the edge cases the old one handled for years.

That is why CodeLead starts with understanding, not rewriting.

Third-party research puts the failure rate of modernization programs between 70 and 79 percent, counting overruns and outright abandonment. Why projects stall

The first step of a modernization

Start with a Legacy System Dossier

A fixed-scope, read-only assessment of one business-critical system. Nothing is modified.

In three to four weeks you receive a practical modernization package. Every section is labeled plainly: what the tool measured, what an engineer judged, what a person confirmed, and what is still an assumption. A receipt for every recommendation.

  • System inventory and architecture map
  • Business rules, each with the line of code behind it
  • Dependency and integration map: the real blast radius
  • Runtime and package lifecycle risks
  • Testability map, and a plan to lock in what must not break
  • Migration-risk register
  • Recommended migration sequence
  • A recommended first move
  • Evidence appendix
  • A living map of the system your team keeps

After the assessment

Then modernize one verified increment at a time.

Once the system is understood, CodeLead helps turn the plan into action. Each step is small, scoped, reviewed, validated, and recorded, and what must not break is locked in with tests before it is changed.

The goal is not a heroic rewrite. It is a sequence of changes the business can trust.

Who decides

The model proposes. CodeLead checks. Humans decide.

The model writes code; it does not decide what is safe. Hard safety checks, tests, human approval and an evidence record control the flow, and people sign off on the goal, the plan, the scope, and anything risky before it lands.

Proof · The agent, unattended

Medium local models can do more than people expect.

Many teams assume there are two options for AI coding: a small local model with weak results, or a frontier agent with the cost, the privacy exposure, and the unpredictable bill.

CodeLead is built on a different bet. A medium-sized model does not need to act like an autonomous genius when the process around it is disciplined: a scoped task, the right context, a plan, hard boundaries, validation at every step, retry rules, and an evidence trail. CodeLead controls the process so the model can punch above its weight.

In the Silicon Exchange case study, a 27-billion-parameter model on one laptop built and verified a complete five-route web application unattended. The same public challenge had needed a 2.8-trillion-parameter model running on four machines.

  • 1h 46mlaunch to exit, unattended
  • 17 / 17increments verified, each gate-checked
  • 30 / 30business-rule oracle cases, checked independently of the model's own tests
  • 0human inputs after the prompt

Not a speed claim: a frontier cloud agent finished the same challenge in fifteen minutes. The claim is that a fixed-cost laptop, with no code leaving it, produced a verified result.

The same discipline is what makes a legacy system safe to touch

  • Understand the code before changing it
  • Keep people in charge of the plan
  • Keep the code on hardware you control
  • Validate every step
  • Leave evidence for the next maintainer

The CodeLead CLI is in private beta. Request early access, or check the receipts first.

Who it's for

Two ways in. One method.

Engineers and teams

You want an agent that works on your code without leaving your machine, that refuses instead of guessing, and that leaves a git history a person can read. The CLI is in private beta; tell us what you would point it at first.

Companies with a legacy system

You need the system understood before anyone touches it, and every change afterwards proven. It starts with a fixed-scope, read-only assessment, and the first call is a diagnosis, not a pitch.

Built in Kirkland, Washington by former Amazon and Microsoft engineers who have lived through large VB6 and C# modernizations. Three U.S. provisional patent applications cover the core methods.

What you keep

You stop guessing. You start deciding.

Whatever the task, CodeLead leaves your team with durable assets. Even if a modernization never goes past the Dossier, you keep these.

  • A map of the architecture and its dependencies
  • A record of the business rules, each with its source
  • A risk-ranked modernization roadmap
  • A validation strategy for every phase
  • A living map of the system your team keeps using
  • Evidence behind every major claim

Tell us what you would point it at.

A code base you want the agent on, or a system you need understood before anyone changes it. We reply within one business day.