V1 · Claude Code first · Codex / Gemini in V2 · MIT

From requirement to tested code, under control.

AIWS (AI Software Factory) drives AI agents through explicit phases: analysis, design with a test specification, small task-by-task implementation with unit tests, then review. Every step is limited to its write scope, every output lives in git, and people approve the decisions that matter.

Workflow

One requirement, nine steps, two human gates

A deterministic orchestrator decides which phase may run; the AI never changes state on its own. People step in at two gates only: approving the design and merging the PR.

  1. 00DiscoverMap the system from source and legacy code
  2. 01AnalysisAcceptance criteria as Given/When/Then
  3. 02Design + Test specDecisions, OpenAPI, test cases from the ACs
  4. GATEHuman approves designApproval is bound to a content hash
  5. 03PlanningSmall tasks, ≤ 10 files, each with its own tests
  6. 04 · loopImplementationDeveloper → build → unit test → commit
  7. 05ReviewChecked against design, contract and standards
  8. GATEHuman merges PROn GitHub or GitLab, as usual
  9. 06Knowledge updateKeeps the map current for the next requirement
Done by AI Human gate Repeats per task, retry ≤ 3, then blocked

Architecture

Three separate layers

Switching AI tools does not mean rewriting the process: adding Codex or Gemini is just another adapter.

Corewhat to do

Agents, skills, workflow, contracts and policies as Markdown/YAML in aiws/, neutral to any AI. Owned and reviewed by people, like code.

Orchestratorwhen it may run

The aiws CLI owns state.yaml, runs phases, validates contracts, enforces diff-scope, runs the real build and tests, and commits with traceability trailers.

Adapterwhich AI does it

claude runs Claude Code headless with limited permissions; scripted runs deterministically for token-free tests. V2: Codex, Gemini and cross-model review.

Enforcement

Gates enforced by real mechanisms, not just prompts

Four independent layers. If one is bypassed, the next one still blocks.

01

Orchestrator

Agents of a locked phase are never called. Approvals are bound to content hashes; retries are capped.

02

Tool permissions

Each headless run gets only the tools it needs; shell access is derived from the project's own build and test commands.

03

Guard hook

PreToolUse blocks out-of-scope writes, secret reads and any attempt by the AI to approve itself, including nested shells.

04

Git + diff-scope

After every run, any file changed outside the scope is reverted and the run fails, however it was written.

Any language

Frontend, backend and legacy in any stack

The orchestrator assumes no stack. aiws detect inspects every source-* folder and proposes build and test commands; add as many sides as you need (fe, be, mobile, worker...).

  • Stack detection: Maven, Gradle, .NET, npm/pnpm/yarn, Python, Go, PHP, Ruby, Rust, Dart/Flutter.
  • Read-only legacy: a language summary (e.g. .cbl, .frm) feeds discovery.
  • Traceable in every framework: TC-3 attached through @DisplayName, [Trait], pytest.mark, t.Run...
Javamaven · gradle Kotlingradle C# / .NETdotnet TypeScriptnpm · pnpm · yarn Pythonpytest Gogo test PHPphpunit Rubyrspec Rustcargo Dart / Flutterflutter test Legacyread-only Other stacksdeclare commands

International standards

Standard output that never breaks your structure

The project's own conventions come first, then the language's official style guide. Renaming, moving or reformatting outside the approved design is reported by the reviewer.

AreaStandard
Acceptance criteriaGiven/When/Then (Gherkin); requirement quality per ISO/IEC/IEEE 29148
Test specificationISO/IEC/IEEE 29119-3; ISTQB test design techniques (partitioning, boundary values, decision tables)
Unit testArrange-Act-Assert; framework naming; TC ids through standard metadata
CodeGoogle Java Style, .NET conventions, PEP 8, Effective Go, PSR-12, Effective Dart…
APIOpenAPI 3.x, RFC 9110, errors as RFC 9457 Problem Details, ISO 8601 timestamps
SecurityOWASP Top 10 / ASVS
CommitConventional Commits with git trailers REQ-ID · Task · Tests · AIWS-Run

Traceability

From requirement to commit, no database needed

Only consistent ids, commit trailers and evidence files in the repository. aiws trace builds the matrix; an AC without a test or a test without a commit blocks the review phase.

# git log -1
feat(REQ-001): T1 BE - setNickname

REQ-ID: REQ-001
Task: T1
Tests: TC-1, TC-2
AIWS-Run: run-0005
ACTest caseTaskCommitResult
AC-1TC-1T1882be59pass
AC-2TC-2T1882be59pass
AC-3TC-3, TC-4, TC-9T1882be59pass
AC-8TC-14T29019a59pass
AC-9TC-15…17T29019a59pass

Real run

Results on a sample requirement with Claude Code

A small frontend + backend requirement run from discovery to knowledge update, headless, with the guard hook enabled.

2/2
tasks passed on the first attempt
11 → 17
ACs traced to test cases, 100%
0
out-of-scope writes
2
human interventions: design approval, PR merge
22
automated tests, CI on Windows, Linux and macOS

Get started

Up and running in a few commands

Requires Node.js ≥ 22, Git and a signed-in Claude Code. Put your code into source-fe/, source-be/ and source-legacy/, then:

  • Human-only commands (approve, reject, answer, resume) refuse to run inside an AI session.
  • Parallel work: several requirements can run analysis and design at once with aiws new --worktree.
  • Token-free tests: the scripted adapter runs the whole workflow deterministically.
# install the CLI
$ cd aiws/adapters/cli && npm install && npm link

# detect stacks, generate Claude config, map the system
$ aiws detect --write
$ aiws sync claude
$ aiws discover

# one requirement
$ aiws new REQ-001
$ aiws run REQ-001            # stops at the design gate
$ aiws approve REQ-001 design
$ aiws run REQ-001            # stops at the PR gate
$ aiws approve REQ-001 pr     # after the merge
$ aiws run REQ-001 && aiws trace REQ-001