webvise ships client software through a documented, AI-powered development workflow: plan mode before any code, a DESIGN.md before any UI, a public skill library, and verifier agents that never approve their own work. The standing proof is webvise.io itself, 100% AI-coded via Claude Code, scoring 96 on Lighthouse, with the repo public at github.com/sebastiankehle.
One receipt up front: a 61-page bilingual website for an architecture client, rebuilt in one day and QA'd with a report covering 60 page loads and zero console errors. You probably have the same tools installed and the output still looks generic. This article documents the pipeline that closes that gap. Six build stages, three model tiers, three agent roles, and the numbers each one has produced.
- Every build runs six named stages: Validate, Scaffold, DESIGN.md, Build, Polish, Ship. The design constraints exist as a file before any UI code does.
- Models are routed in three tiers. The strongest model plans, an efficient model implements, a cheap model does lookups. The expensive model authors a skill once and cheap models run it from then on.
- A verifier agent never self-approves. Review runs in a fresh context with no memory of writing the code, and UI changes get exercised in a real browser before shipping.
- Nine skills live in the public repo, from animation-vocabulary to shadcn. Registering a new one is a single symlink line.
- The receipts are public: webvise.io is 100% AI-coded with a 96 Lighthouse score, and the same pipeline has delivered 25+ client projects at a 3-week average to launch.
The pipeline stays the same whether the deliverable is a marketing site or a portal with auth, roles, and dashboards. If you need that class of software shipped, webvise's full-stack application service runs every engagement through the stages below.
The six-stage build sequence
Every project moves through six stages in a fixed order, and each stage has an exit condition that gets checked before the next one starts. The sequence is boring on purpose. Boring sequences survive deadline pressure.
| Stage | What happens | Exit condition |
|---|---|---|
| 1. Validate | Scope, content, and success criteria get pinned down | A build decision in writing |
| 2. Scaffold | Better-T-Stack generates the base, the example app gets deleted | A clean repo that builds and runs |
| 3. DESIGN.md | Colors, type scale, spacing, and motion rules written as a file | Constraints the model must obey |
| 4. Build | Features implemented against the plan, plan mode before code | Working software behind every route |
| 5. Polish | Motion, empty states, edge cases, copy | Review by a separate verifier agent |
| 6. Ship | Browser-verified QA, then deploy | Zero console errors across all pages |
Stage 3 is the one most teams skip, and skipping it is what makes AI-generated interfaces look generic: the model has no constraint to obey. A DESIGN.md pins the palette, the type scale, the spacing system, and the motion rules before any UI code exists. "Make it look good" becomes a set of checkable instructions.
Stage 2 carries one habit worth stealing: right after Better-T-Stack scaffolds the project, its example app gets deleted. Demo code that survives into week two becomes load-bearing.
Three model tiers and one economic rule
The strongest model plans. An efficient mid-tier model implements against that plan. A cheap model handles lookups, doc searches, and file reads. Routing this way reserves the expensive context for the decisions that branch the project and keeps costs low through the mechanical middle.
One rule carries the economics: the expensive model authors a skill once, then cheap models execute it on every later run. Authoring is the judgment-heavy part, execution follows patterns, and the price of each should match.
Orchestrator, executor, verifier
Agent work splits into three roles. An orchestrator holds the plan and dispatches tasks. Executors implement one task each, with only the context that task needs. A verifier checks the result, and it never runs in the same context that produced the code.
Self-approval is the failure mode this prevents. A model reviewing its own output inherits its own assumptions, so the review confirms the code instead of testing it. A fresh verifier context has to reconstruct the reasoning from the diff, and the gaps show up exactly there.
For UI work the verifier's job is physical: run the dev server and exercise the change in a real browser. The 61-page bilingual rebuild for Gabrys Architektur went through that gate in one day, and the handover included a QA report covering 60 page loads with zero console errors.
A skill library instead of prompts in chat history
Recurring judgment gets packaged as skills, instruction files the agent loads when a task matches. Nine live in the public webvise repo: animation-vocabulary, apple-design, blog-article, design-webvise-asset, emil-design-eng, find-animation-opportunities, improve-animations, review-animations, and shadcn. Registering a new one for Claude Code is one symlink line.
The same mechanism ships to clients. MP Bau, a construction client, got a document-conversion skill packaged as a client-facing tool. For Gabrys Architektur, three agents wrote blog content in parallel while sharing one authoring skill, so every article came out in the same voice.
The MP Bau build has its own writeup in the document-automation article. If you are still picking which workflow in your company deserves a skill first, webvise's AI consulting starts every engagement exactly there.
AGENTS.md is the contract
Every repo carries an AGENTS.md the agent reads before touching a file. It holds the command table (dev, build, lint, test, db), the conventions (Biome for formatting, typed env vars, Server Components by default), and the behavioral rules that keep output predictable.
- Read before editing. No change lands in a file the agent has never opened.
- Minimum code that solves the problem. No speculative abstractions, no configurability nobody asked for.
- For UI work: run the dev server and exercise the change in a browser. The rule that turns "looks done" into "verified done".
A production-ready template with the full rule set is in the AGENTS.md template article. Copying it takes five minutes and removes a whole category of agent misfires.
The proof is public
webvise.io is the pipeline's standing demo: 100% AI-coded via Claude Code, 96 on Lighthouse, repo public at github.com/sebastiankehle. Every page, every animation, and the blog system itself went through the six stages above.
| Project | Shipped through the pipeline | Numbers |
|---|---|---|
| webvise.io | Marketing site, blog, dashboard | 100% AI-coded, 96 Lighthouse |
| Gabrys Architektur | 61-page bilingual site rebuild | 1 day, 60 page loads QA'd, 0 console errors |
| MP Bau | Document-conversion skill as a client tool | Authored once, runs on cheap models |
| All projects | Client software across 25+ deliveries | 3-week average to launch |
Every claim above has a public URL or a repo behind it. If your next build should run through this pipeline, get in touch with webvise and I will map the six stages onto your project.
Development practices are aligned with ISO 27001 and ISO 42001 standards.