Status, Specs & Roadmap

Feature-to-completion mapping against the SOW and PR/FAQ · spec quality & coverage · roadmap. Every status is read from the repo, not estimated.

Current status snapshotAs of 2026-09-06, reconciled against the on-disk specs and the GitHub tracker.

40%
backlog completion (27.5 / 68 items)
13 / 39
specs implemented (10 partial · 16 not)
853 / 2,525
spec task boxes ticked (34%)
3
designs proven RTL→GDSII
One distinction runs through everything here: "fixed in code" is not "live in the account". Several finished specs ship a feature switched off in infrastructure, and the Well-Architected remediation is 38/38 in the repo but 0 deployed. The security page carries the deployment-state detail.

Two deployed environments, different builds

EnvModeWhat runsProven
POC (primary account)tools in-image, S3 workspace7 stacks UPDATE_COMPLETE · 13 runtimes READY · Well-Architected remediation livecalc4 prompt→GDSII, 7/7 gates PASS
DEV (aes-dev stack)S3 artifact workspace, PCS/Slurm P&R on FSx8 stacks · 13 runtimes · 3 PCS compute groups ACTIVE (min0/max1)AES-128 & PicoRV32 to accepted GDS

SOW deliverables → completionThe 19 in-scope deliverables (D1–D19) from the Statement of Work, mapped to what the code actually does today.

#DeliverableStatusEvidence / note
D1Step Functions pipeline DAGDelivered87-state ASL owns the full RTL→GDSII flow; shipped pipeline spec 109/109
D213 agent runtimesDeliveredOne arm64 image, 13 roles; lint renamed verification (also hosts formal)
D38 deterministic gate LambdasDeliveredNow 13 gate Lambdas / 7 verdicts on a normal run; no model SDK importable in gates
D4Async task-token integrationDeliveredwaitForTaskToken + 60s heartbeat, exactly-one signal on all 13 agents
D5PCS/Slurm + FSx LustrePartialP&R submits to Slurm (live on DEV, ACT-96 FSx workspace). Simulation still runs in-image everywhere
D6Per-agent IAM rolesDelivered13 scoped roles; none holds InvokeAgentRuntime
D7Tool-agnostic adapter contractDelivered6 OSS adapters on one contract, TOOL_REGISTRY-driven; commercial tiers declared mocks
D8Bounded fix loopsDeliveredState-machine counter, schema + iverilog compile-check before re-entry; run-wide + per-stage budgets
D9Formal verification gateDeliveredSymbiYosys + z3 deployed & proven live (2026-08-12); vacuity census demotes zero-assertion proofs
D10AgentCore Memory (semantic recall)PartialCode complete + dedup guard; MEMORY_ID reaches no runtime — dead config today
D11RAG pipeline (OpenSearch + Titan)Not builtNo RAG index in the shipped system
D12MCP Gateway + CognitoPartialGateway/eda_adapter + Cognito exist but reachable only from the legacy entrypoint; no production caller
D13Observability pipelinePartialTelemetry→Glue→Athena→QuickSight live; GenAI tracing (OTEL) built on integration, not deployed
D14CodeBuild image factoryDeliveredARM_CONTAINER two-image build → ECR; base hash guards against stale-base builds
D15CloudFormation IaCDeliveredLayered stacks foundation→platform→runtimes + compute/notifications/quicksight/cost/gui
D16Run registry + notificationsDeliveredDynamoDB run record with cost/token rollups; SES/SNS + presigned artifact bundle
D17Acceptance tests (calc4, AES-128, PicoRV32)PartialAll three proven RTL→GDSII; AES-128/PicoRV32 as measured runs, but no automated multi-design runner
D18DocumentationDeliveredArchitecture, setup, runbook, prerequisites, benchmarks, backlog — 50 docs
D19Optional human sign-off gateSpec onlypipeline-human-release-preauth spec exists (0/86); H0/H1/H2 states absent
Two acceptance criteria describe a design the measured runs do not match. AC-7 was written as AES-128 at ~36K instances; the core that runs is 9,710 cells (iterative 11-cycle, ~36K is the unrolled scale). AC-8 was written RV32IMC; the measured core is RV32IM from a frozen package. Amending an acceptance criterion is the owner's call — flagged, not silently reconciled.

The PR/FAQ narrativeThe problem, the promise, and the honesty the FAQ insists on.

The problem

Semiconductor front-end design is 1–3 days of senior-engineer time per standard block, most of it repetitive rule-governed work — and the talent shortage means teams cannot hire fast enough. The routine 60–70% (FIFOs, bus interfaces, encryption cores, state machines, ALUs) follows well-understood patterns, but existing AI copilots offer only code completion, not end-to-end verified design.

The promise

"The critical innovation is the trust model. Every pass/fail verdict is made by a deterministic Lambda function that parses tool exit codes and log files — no AI is involved in any verdict decision. The AI writes the code; mathematical solvers and EDA tools judge it."

Plain-English spec → validated routed GDSII in 3–15 minutes, running entirely in the customer's AWS account with zero design-IP egress, at ~$0.67–2 per run (measured on calc4).

What the FAQ is careful about

Spec inventory — quality & coverageEvery .kiro/specs/ folder mapped to implementation status, re-derived from disk. A checked task box is not a live behaviour — the Evidence column names code or a test.

Implemented 13 · Partial 10 · Not built 16  — 39 rows
33% of specs finished. The finished specs are the small ones — 34% of all planned task boxes are ticked.
SpecTasksStatus
pipeline-chip-design-orchestrator109/109Implemented
agents-agentic-optimization96/96Implemented
eda-tool-exit-code-fidelity95/95Implemented
gui-desktop-client104/104Implemented
eda-waveform-visualization63/63Implemented (not yet in a deployed image)
agents-notification-envelope41/41Implemented (pending deploy)
agents-runtime-session-identity10/17Implemented (code-complete, deploy-unverified)
obs-agent-tracing63/63Implemented (on integration, pending deploy)
gui-optional-stage-selection19/19Implemented
pipeline-optional-stage-selection24/24Implemented (CLI half)
pipeline-mpw-precheckImplemented
pipeline-prelayout-staImplemented
gui-gds3d-qt-test-allocation5/6Implemented
agents-bedrock-guardrails13/26Agent code complete, provably inert
agents-memory-strategy-upgrade17/28Code complete, MEMORY_ID unwired
gds-3d-viewer40/99Builder + auto-load works, rest is a shell
gui-chat-style-run-workspace0/65Mostly implemented, acceptance unrecorded
obs-business-metrics-fidelity21/73Partial
ops-cost-optimization1/37Partial (thread config + Graviton5 shipped)
ops-eda-plane-separation0/71Partial (PCS P&R + FSx landed; sim boundary reserved)
pipeline-frozen-reference-package8/9Backend + CLI implemented
pipeline-signoff-gate-spineSubstantially implemented
gui-run-history-trackingplaceholderConfig-only placeholder
agents-complex-design-generation0/18Not built
agents-eval-harness0/74Not built
agents-model-config-automation0/102Not built
agents-taxonomy-realignment0/20Not built
eda-physical-signoff0/37Not built
gui-concurrent-run-submission0/95Not built
gui-run-start-end-time0/66Not built
obs-dashboards-agent-analytics0/91Not built
obs-visualization0/72Not built
ops-acceptance-automation0/82Not built
ops-ci-cd-pipeline0/84Not built
ops-fsx-workspace-lifecycle1/59Not built
ops-security-acceptance-evidence0/86Not built
pipeline-coverage-measurement0/79Not built
pipeline-human-release-preauth0/86Not built
pipeline-repair-taxonomy-corpus0/90Not built
Spec quality is high where it counts. Every spec carries requirements/design/tasks; the inventory is reconciled by a script (no hand-maintained counts), and "not built" specs are genuine plans with acceptance criteria — not stubs. Known hygiene items: 3 folders were unrowed and 1 was double-rowed at prior passes (now caught by a 4-check reconciliation), and 3 specId values collide across 6 .config.kiro files (harmless to text, flagged for tooling).

Backlog completion68 numbered items ever opened; completion counted by script, not by hand.

40% = (21 archived + 5 closed-in-place + 0.5×3 partially-closed) / 68. "Partially delivered" items contribute 0 to the numerator by design, so a landed half is visible without inflating the figure.

21
closed & archived
5
closed in place
9
partially delivered / closed
33
fully open

Highest-value open work (P1)

Three other tallies, none the same number: spec tasks 853/2,525 (34%) · Well-Architected remediation 38/38 in repo, 0 deployed · tape-out readiness 1/30 tasks formally, Phase A substantially built.

Roadmap — tape-out readinessWhat stands between the GDSII candidate and a foundry-accepted handoff. Nothing below is broken; it is absent, and the repo says so consistently.

Phase A — physical implementation completeness substantially built

Real SDC timing constraints (T1), clock tree synthesis (T2, landed), power grid (T3, pinned PDN landed), filler/tap cells (T4), timing repair (T5). Running DRC/LVS before these produces artifacts, not real defects — which is why this phase is first.

Phase B — physical verification gates not built

Magic DRC + Netgen LVS adapters and gates (comparator settled as Netgen), antenna (done) + density, route-DRC promoted to a verdict, ASL wiring, manifest schema v2. The single highest-leverage unlock: the post-GDS sign-off spine is deployed but returns required:false for every prompt-generated design.

Phase C — netlist-level & timing verification not built

Logical equivalence RTL↔netlist↔P&R (no equivalence check exists at any link today), gate-level sim with SDF, extraction + multi-corner STA on the worst corner, power/IR drop, CDC.

Phase D — foundry handoff not built

NDA PDK handling (on FSx, never in the image), foundry signoff decks via the Gateway http/batch tier, chip-level structures (pad frame, seal ring), ECO flow, immutable handoff bundle with Object Lock.

Adjacent roadmap themes