Automation

From idea to working app in 8 stages, with a check on every task

A sequence that takes a one-sentence idea to a finished app with a person involved at only two points: how it works, the rules that hold it together, and what the numbers from the first app do and don't show.

Pixedi AI Lab
  • 6 min read

SummaryUnder a minute

The short version

The test: whether a single sentence could become a working app with a person speaking only at the start and the end. The answer depended on three rules: every task carries a machine check, the coding model never writes its own check, and each model is asked only for what it can deliver. The first app passed all its checks and tests; a person still made the final call.

Key takeaways

  1. An 8-stage sequence has a person give the idea at the start and judge the finished app at the end, and everything in between is automatic.
  2. Every automated task must carry a machine check where exit code 0 means pass, and a task without one is rejected at the task-split stage.
  3. The model that writes the code never writes or executes its own check, because a model grading its own work can loosen the test until it passes.
  4. Models that can't write to disk return file blocks that the system writes for them, which let a local model write 2 files in 33 seconds over 2 turns and pass its tests.
  5. On the first app, all 289 checks passed and 115 of 115 app tests passed, which describes one app and predicts nothing about the next.
A tidy office desk in afternoon light with a notebook sketch of connected boxes beside a closed laptop.

The question

The test is whether a one-sentence idea can be carried all the way to a finished app with a person involved at only two points, with a result that can still be trusted. The reason this matters is a specific failure the whole design works around: AI models quietly produce code that looks like it works and doesn't, and without a check that a machine actually measures, that kind of mistake stays hidden.

The design

The sequence has 8 fixed stages. A person speaks at two points only, at the start with the idea and at the end with a verdict on whether it works, and everything in between happens automatically in this order:

  1. Idea: a person writes one sentence describing the app, and that sentence is all the sequence starts from.
  2. Questions: a planning model turns that sentence into research questions and saves them as a structured list the next stage can read.
  3. Research: a second model works through each question and writes up what it found along with where it found it, so later stages can point back to the sources.
  4. Architecture: the planning model writes a project description and breaks the app into components.
  5. Task split: the planning model splits the work into tasks, assigns each one to a machine and a model, and writes an independent check for every task.
  6. Build: the models that write code each pick up their own share of the tasks and work through them.
  7. Machine check: plain code, with no model involved, executes every check and records the result for each task.
  8. Human review: a person looks at the finished app and answers one question, "does it work?"
A worn oak workbench in a bright workshop with a laptop, a rolled-up paper plan, a pencil and a mug.

Operating it comes down to a handful of short commands: one opens a project from the idea, one moves it forward a single stage, one carries it automatically up to the human review, and one lists every project with its current stage. Two more either approve the result or send it back with a note on what to fix.

Behind the sequence sit five models reached through one interface: a hosted planning model, a research model used through a chat subscription, two hosted coding models with usage quotas, and a free local model. The sequence never asks for a model by name. It asks for a kind of work (planning, research, hard coding, high-volume coding), and each kind has its own order of fallbacks. Hard coding goes to the quota-limited models first, so those tasks are kept few and small, and high-volume coding goes to the local model first because calling it is free. Research moved to the chat model first after its plan stopped counting plain-text use on August 6, 2026, so the token-heavy reading happens there and the planning model gets only a small number of high-value calls.

The whole thing depends on three rules, and each one closes a way the system could fail without anyone seeing it.

Every automated task carries a machine check

When the planning model splits the work, each task has to come with a check: a command that executes in a terminal, where exit code 0 means the task passed and anything else means it failed. A task split that contains a task without one is rejected and has to be redone. Models produce code that seems to work and doesn't, and they do it quietly, so the only reliable signal is a check a machine can measure.

Some work can't be measured that way, like opening an account, setting up a payment or filling in an app store submission form. Those tasks are marked as human work and go to a person. That is the honest limit of the system, and marking it beats pretending.

A printed checklist on a desk with a tick beside every line.

The engine that writes the code never writes its own check

Having plain code execute the checks isn't enough on its own, because if the model writing the code also writes the check file, it can loosen the check until its own work passes. So the planning model writes each check as a separate small test script, one per task, kept in its own folder that the coding model can't touch. A model approving its own work is the point where a lot of automatic coding setups fail unnoticed.

Instructions match what each model can do

Three of the five models can write files directly into a working folder. The local model and the research model can only send back text, and so will any hosted model connected later through an API. Give a text-only model the instruction "write these files" and it sends back text, the check reports that the file doesn't exist, and the task drops out without a visible error. So models that can't write to disk are asked for their output as file blocks (each file's name followed by its contents), and the system writes those blocks to disk itself. The first question before adding any new model is whether it can write its own files into the working folder.

Measurement

The yardstick is the exit code. A task counts as done only when its check exits with 0, and the app also has its own test suite, which has to pass in full.

This has limits worth stating plainly. A check only proves what it tests, so a weak check will happily pass weak code. The checks are written by a model too, just a different one from the model that wrote the code, which removes the conflict of interest without making the checks perfect. Tasks marked as human work aren't covered by any of the machine numbers, and these numbers come from the first app, so there is no basis yet for saying what other apps would look like.

That is why the last stage is a person. The machine checks show that the parts behave the way the planning model said they should, and the person decides whether the app does what the idea asked for.

Results

The file-block approach was the first thing to confirm, because without it the local model couldn't contribute at all. Given block-style instructions, the local model wrote 2 files in 33 seconds over 2 turns, the system wrote them to disk, and they passed their tests.

On the first app built end to end, all 289 checks were green (passed) and 115 of 115 app tests passed, and then it went to a person for the final question.

Those numbers describe one app. They show the sequence can hold together on a real build, and they say nothing about what the next app will score.

A person at a standing desk reviewing something on a tablet held away from the camera.

Takeaways

Make the check a condition for the task

Writing tests after the code is easy to skip. Refusing to accept a task without a check moves testing to the front, where the planning model has to define what "done" means before any code exists.

Keep the writer and the grader apart

Separating who executes the check from who writes the code isn't enough on its own, because the file that defines the check also has to be out of reach of the model doing the work.

Ask each model only for what it can deliver

The quiet failures came from mismatched instructions, and the fix was cheap. Asking text-only models for file blocks is what let the free local model take on real coding tasks.

Be honest about what goes to a person

Marking unmeasurable tasks as human work keeps the machine numbers meaningful. If those tasks slipped through with a fake check, every green result would be worth less.

Two things deserve more attention next time: reviewing the checks themselves early on, since the whole system leans on them, and putting several more apps through the sequence before drawing any conclusion about how well it holds up in general.

Ask AI about this AI Lab note

Opens the assistant in a new tab with this page as the source.

Keep reading

Want this handled for your business?

Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.