Case studies
Ten systems on one shelf, each told the same way: the problem, what I built, and what changed. The work studies are sanitised (no client names, no internals); the personal ones are fully mine, so they can be shown in detail.
- A The composite metric 60 of 171 clients, 60% of volume The one number that kept the biggest client An industry-first composite metric, born from years of sitting with the clients who needed it, that fused AI sentiment analysis with operational ratings into a single score operators could run a business on. retained the client worth 50%+ of revenue
- B The backlog teammate 4 epics in parallel, one dev An AI teammate that works the ticket backlog A three-skill workflow that specs, decomposes, and orchestrates a team's Jira backlog, letting one developer run several epics in parallel while agents work the well-specified tickets. one epic per fortnight became four epics in parallel
- C The self-testing agent 0 humans in the verify loop An AI that tests its own work in a live cloud environment A tester agent that syncs finished work into an isolated cloud environment, executes it in the real runtime, and reacts only to structured, machine-readable results. edit-to-verified loop with no human in the middle
- D The 250-branch engine £30k+ a year, automated Replacing a full-time reporting role for 250 branches When the biggest client made replacing a manual reporting job a condition of renewal, a paginated-report engine took over the work, running seven days a week instead of five. a £30,000+ a year manual process, automated and expanded
- E The self-reporting pipeline every night reports its own runs AI-driven reliability engineering for an overnight data platform Months of run history audited by agents into a defensible failure-mode taxonomy, then a self-observation layer so the platform reports its own nights. from "we find out when a client complains" to a daily evidence-based digest
- F The reporting platform 400 reports, no hands The reporting platform: 400 branded reports without hands A nightly notebook engine that generates, quality-checks, and delivers hundreds of branded client reports, plus a generative skill that builds new Power BI reports from a one-line prompt. 30 to 80 hours of weekly manual work replaced by a run that takes minutes
- G The model fleet 171 of 171 models, first run Fleet operations: 170+ semantic models, refresh, rebinding, and cost control One overloaded all-client model became 171 per-client models maintained from a single baseline, ending daily capacity upgrades and putting cloud spend under active management. near-daily capacity upgrades stopped the week it shipped
- H Courting 3 platforms shipped solo Courting: a three-platform product shipped solo, with agents doing the shipping A racquet sports management platform on iOS, Android, and web, delivered through a 14-skill agentic pipeline that specifies, implements, tests, and releases. live product; 2,100+ tickets and 2,000+ PRs with agents in the loop
- I MindGrid 2nd stack same discipline MindGrid: a consumer mobile game, AI-assisted, off the Microsoft stack A React Native word-puzzle game with a Supabase backend, a tiered test harness, and its own internal level-generator tool. shipped consumer software on a second stack
- J The websites 3 sites live, brand to deploy The websites: design and build, end to end, with AI Three live sites, including the one you are reading, designed and shipped with AI across the whole craft: brand identity, art direction, generated imagery, working code, and browser-verified quality. three live sites, brand to deploy
At work · a hospitality feedback analytics SaaS company On my own time
Pull a book. Click to open the full study.