One person, shipping like a team.
The products aren't the hard part. The harness underneath them is: guardrails that keep it safe, governance that keeps it honest, and enough flexibility that one business person can build, test, and adapt real software with no developer and near-zero cost. Aaron Patzalek, 15+ years in category development, innovation, and customer experience, started Two Birds Innovation in early 2026 and built that system first. This page shows what's running on it.
A software studio that runs itself overnight
Most people using AI are using it like an assistant. This is using it as the engineer. Months went into building the infrastructure, the shop floor, before worrying about the products that sit on top of it.
An idea can land in a backlog at midnight and have committed code, passing smoke tests, and a health report waiting by morning. It does not even have to be typed: hold a key and talk, or send a voice note from a phone. Speech is transcribed on the machine with Whisper, not by a cloud transcription service. One person running a studio that ships like a five-person team.
The cost structure: $0 hosting, $0 npm packages in any product repo. Sovereignty-first, and as close to free as possible. Nothing runs on pay-per-use AI. The work rides subscriptions already paid for (Claude, ChatGPT, Google AI Pro and OpenCode Go), spread across four engines so no single vendor is a single point of failure, and the one unattended route that used prepaid API credit is switched off. There are a small number of named, bounded exceptions (currently a low-cost AI video tool), chosen because the output stays owned and portable, not because sovereignty stopped mattering. No vendor can price you out of the system you built.
The stack
Counted from the repository and the live configuration. The figures below are recomputed by the nightly build, and the date on each one says when it last was.
overnight loop
(incl. reviewed vendored packs)
the engine configs
for specific sprint types
Engines (who does the work):
MCP servers (how engines reach tools):
Voice and phone:
Automation, and tools being weighed:
Also available on demand in Claude.ai (account connectors, not counted above):
📖 Stack 101: the tools in the stack, in plain English →
What each tool does, whether it is in or out, and why.
From idea to live — one sprint
Same flow, grouped by what each stage does: an idea is captured, routed to an engine, executed in isolation, and verified before anyone calls it done.
Hold Ctrl+Space and talk, or send a Telegram voice note from a phone. Dictation and Telegram voice notes are transcribed on the machine with Whisper, not by a cloud transcription service. Forwarded videos and links are transcribed, scored and filed the same way.
The idea becomes a backlog item with a priority, an effort and return rating, and an owner. An autonomy check blocks anything an agent can do itself, so only work that truly needs a person reaches the human queue.
An operating-system timer, not an open terminal, wakes the supervisor every ten minutes. It picks the next job by priority, checks it for collisions and design gates, and sends it to an engine that has headroom.
Claude Code, Codex, OpenCode and Antigravity share the work. The two main paid accounts have a weekly usage cap, so neither gets drained. A capped, down or switched-off engine is skipped automatically and the next one takes the job.
The engine works in a throwaway clone and never touches the live working tree. It may only edit the paths the job declares, it cannot delete or rename files, and its result is secret-scanned before it can land.
No build step, no Node.js, no npm in the product repos. Git push, and GitHub deploys in seconds. $0 hosting. WCAG accessibility checks run automatically via GitHub Actions on every push.
Tests run first. For anything on a live site, the real page then has to pass a Playwright check. Landing is fast-forward only, so a failed job cannot overwrite good work.
At 2am the overnight build re-checks every product, re-syncs the repos and recomputes the numbers on this page. A health check runs whenever a session starts, and scheduled jobs leave a heartbeat, so a job that goes quiet is flagged within a day or two.
Every step above runs without a person. This is the one node that is not automatic: the decisions, approvals and judgment calls that only a human can make. Everything an agent can resolve on its own never reaches it.
Five layers. Built to run without babysitting.
Products on top. Automation and orchestration in the middle. Several AI engines, held to hard rules, as the workforce. Sovereignty principles at the base. Every choice above L5 must survive the test: "if this vendor disappeared tomorrow, what breaks?" (The name's a nod to 2001: A Space Odyssey. Make of that what you will.)
The hard rules — enforced on every sprint
These aren't preferences. They're gates. If a rule fires, the sprint stops until it's cleared.
Four conditions required: PRODUCT.md exists, /impeccable audit has run this quarter, a human approved the shape brief for structural changes, dark mode tested on Android Chrome. Any one missing — sprint doesn't start.
Before results commit, the sprint is reviewed by a panel of AI personas: the Scrappy Pack, the Founding Board, and the Inner Circle. A REWORK verdict blocks the commit. The panel catches what a solo operator misses when heads-down.
Every email, landing page, or grant submission is scanned against a banned word list. A compliance tag is appended proving the scan ran. Nothing goes out unscanned.
Before any SaaS, API, or dependency lands, the decapitation checklist runs, and the tool's written rationale is read first: why it was chosen, what it replaced, and what a swap would have to beat. If this service disappeared tomorrow, what breaks? No paid service without proof no sovereign alternative exists.
Three checks: Can PowerShell or Python execute this? Can an MCP tool handle it? Does it only need files and scripts? If any yes — the agent does it. Only genuinely human tasks reach the queue.
The guard blocks an identical call repeated three times in a row and any touch of credential files, and warns and logs when a single turn passes 120 and then 200 tool calls. The warnings never stop the work; they make a runaway loop visible.
Nothing that touches a live site is done until a Playwright check passes against the real URL. Domain and DNS changes also get a check through the real public path, on a deep page, not just the home page. Reporting "done" past a failing check is treated as a serious defect.
Deleting data, rewriting git history, or removing cloud resources needs an explicit yes from a human in that session. A hook enforces it, so an agent cannot talk its way past it.
The two main paid AI accounts each have a weekly usage cap of 75%. Past it, the engine is skipped until its weekly reset, so no subscription is run dry and no engine becomes a hidden single point of failure.
A new scheduled job is not done when it runs once. It needs a trigger confirmed on the live scheduler, a heartbeat, an alert threshold, and a test where the trigger is removed and the health check turns red.
What runs at 2am every night, and all day
The 11 active product and portfolio repos synced from remote and pushed to GitHub. Fewer than the 15 built, because archived, template and utility repos are not in the nightly set.
Performance, Accessibility, Best Practices, SEO scored. Results written to quality/lighthouse-results/.
Headless Chromium hits all products. Checks key elements, exercises user flows. Screenshots on failure.
Scans every open backlog item. Anything the AI can handle gets flagged for the agent. It never reaches a human.
Emails approved in the outbox folder are sent via smtplib. No SendGrid, no subscription.
Reads session wins and Notion P1 items, writes a morning briefing. Ready when the laptop opens.
Pulls Statistics Canada Table 14-10-0287-01. Extracts national unemployment rate. Updates Career Coach automatically.
The claims check re-counts commits, repositories, skills, engines and products from their sources and stamps the date it ran. If this job does not run, the date visibly stops moving.
An operating-system timer wakes the supervisor. It dispatches the next eligible job, watches the one in flight, and moves a stuck job to another engine. No terminal has to stay open.
Checks the Telegram inbox for forwarded videos, links and voice notes, transcribes them on the machine, scores them against current work and files them.
Scheduled jobs write a heartbeat as they finish. A missing or stale one raises a flag at the next health check, so a job that quietly stops is noticed within a day or two.
What's running on the shop floor
29-module digital literacy platform. Bilingual EN/FR. WCAG AA. Built for the gap left when federal Digital Literacy Exchange Program funding ended on 31 March 2025 (a program that had reached over 650,000 participants). Licensing for libraries and municipalities; pricing is on the DCC for libraries page.
Digital literacy curriculum for children. Online safety, AI literacy, technology judgment. Same trust-first architecture as the adult platform. Built on the same codebase — not a separate product, a major line extension.
AI-assisted job search. Pulls live StatCan unemployment data monthly. CV customization, salary research protocol. B2B model targets employment agencies and college career centres.
Free 15-minute AI readiness diagnostic. Seven questions. Personalized SWOT and action plan. Qualifies consulting leads before they book a call. Every local SME in Elgin County is a target.
Apartment search and tracking dashboard. Started as a personal civic tool for a friend in London, ON. The repository is private and there is no public site, but the daily listing-refresh automation runs against it and has for months. Being positioned as a white-label housing navigator.
A technology-readiness diagnostic with sourced Canadian benchmarks, so an adviser can see where their own practice stands before advising clients on technology. Spec complete; deliberately held at the gate until one positioning decision is made.
Strategy and operations work for businesses that are growing, stuck, or looking for new perspectives. Technology as one tool in the kit, not the pitch. Direct engagement model — you work with Aaron, not a team of juniors. Reach out via the consulting page to explore fit.
Sovereign by design — every tool gets audited
Before any tool lands in the stack, it passes a sovereignty check. The question: if this service disappeared tomorrow, what breaks?
| Function | What we use | What we rejected |
|---|---|---|
| Hosting | GitHub Pages | Vercel / Netlify |
| Scripting | Python stdlib | npm packages |
| Scheduler | Task Scheduler | Cloud cron SaaS |
| Monitoring | uptime-monitor.py | Datadog / PagerDuty |
| Smoke tests | Playwright OSS | BrowserStack |
| Email send | smtplib (Gmail) | SendGrid / Mailchimp |
| Backup mirror | Codeberg.org | GitHub only |
| Frontend | Static HTML/CSS | React / Next.js |
| Model routing | OmniRoute (self-hosted) | OpenRouter / LiteLLM |
What's being built next
A live snapshot of what's in motion right now — some locked in, some still in design. The shop floor keeps adding capability.
A near-zero-cost pipeline that turns a script into a narrated, animated explainer video — bilingual EN/FR. First outputs already rendered. Brings video production in-house without a studio or a subscription stack.
Forward a video or clip and the system pulls the transcript, scores it against what Two Birds is actually working on, routes the strong ones, and catalogues the rest for the record. Built from tools already in the stack.
An invite-based beta so real seniors and the family members supporting them shape the Digital Confidence Centre before public release. Feedback loop and invite flow drafted.
Folding the separate dashboards into a single command deck — every product's status, backlog, and health on one screen, reachable from the phone.
Routes coding work across Claude Code, Codex, OpenCode and Antigravity by real-time capacity and weekly usage caps, so work keeps moving when one tool hits a limit. The supervisor, the disposable-copy runner and the landing checks are live: an overnight run is started on request, then runs without anyone watching. What is still being built is the safety net around it: canary jobs, clearer failure alerts, and engine coverage that is the same everywhere.
Adopting vetted, open brand design systems so any new page can be scaffolded on-brand in one shot — each one cross-checked against the stack's sovereignty rules before it's pulled in.
Send a voice note or a text from a phone and it acts: the audio is transcribed on the machine, then turned into a calendar invite, a formatted Google Doc or a Notion page. Transcription happens on the machine, it runs on subscriptions already paid for, and a safe hold replaces any pay-per-use fallback. Next: voicemail triage and a morning briefing.
Hold Ctrl+Space and talk, with live text while you speak, in any coding tool. The live preview and local transcription work today. Delivering the final text into every tool the same way is being hardened.
Planned: the Hermes Agent trial, once its gate is cleared. Possibly Orca, for running several tools side by side. Already done: the overnight loop moved to the always-on machine (m73), and the Antigravity engine now runs through its command-line print mode on the Google plan, with its pay-per-use API route off.