AI consulting · Production reliability
I take AI systems from demo to production.
Most AI pilots stall between proof-of-concept and P&L. I find exactly what will fail, and hand your team the plan to fix it. Fixed fee, five days.
Consulting experience with teams at
- Microsoft
- AWS
- IBM
- NTT
- Capgemini
- Insight
- Rackspace
- Hewlett Packard
The problem
Your pilot demos well. That was never the hard part.
The cause is rarely the model. It is the gap between a demo that works and a system you can trust when nobody is watching it: retries that silently don’t fire, extraction that drops records without an error, an entry point with no alarm on it. Closing that gap is engineering discipline, not AI hype, and it is the shortest path from generative-AI pilot to P&L.
What I bring
The skillset this problem actually needs.
Depth
PhD in Computer Science; Master’s and Bachelor’s in telecommunications and electronic engineering. I read systems to first principles, not to a checklist.
Translation
15+ years consulting, AI lead on medium-to-large programmes. Findings written for engineers, defensible in the boardroom.
Currency
I build and run production agentic-AI systems on the major clouds every day. This isn’t advisory from someone who stopped shipping.
Agentic AI systems
Multi-step LLM workflows and retrieval pipelines that hold up under real traffic.
Document AI at scale
Extraction that treats messy documents as the normal case, and never drops data silently.
Evaluation, not vibes
Deterministic software you unit-test. Non-deterministic systems you evaluate. I design the evals, baselines and regression gates that turn “is it good enough?” into a measured answer.
Observability that works
Alarms on the paths that serve traffic, and a real answer to “would we know if this broke?”
Production readiness
Retries, idempotency, concurrency, load: the discipline that decides whether AI ships.
Before LLMs, there was ML
Traditional ML, image and video processing, OCR, IoT and robotics: the years of shipped systems that tell me which of your problems actually needs an LLM.
Offerings
Start small. Scale only if it earns it.
Production-Readiness Review
AI readiness assessment · 5 days · fixed fee
One AI system, independently reviewed: prioritised findings, failure-mode and observability analysis, and a remediation plan your engineers can execute.
Request a reviewRemediation Sprint
2–4 weeks · quoted upfront
I work inside your team to close the review’s highest-risk findings: shipped, tested and monitored, not just recommended.
Fractional AI Reliability
Monthly · scoped to your systems
A standing eye on what you’ve shipped: reviews of new features, incident support, and someone who already knows your stack.
How to think about the cost
- Pricing is discussed after a conversation about your project, then fixed and quoted upfront. No meter running, no scope creep.
- A stalled pilot costs a quarter of team salary, and the credibility of the next one.
- If the review finds nothing serious, that’s the cheapest good news you’ll buy all year.
Independent
No tooling to sell, no partner kickbacks. The findings are the product.
Evidence over opinion
Every finding is reproducible: a trace, a test, or a failing input.
Yours to keep
Reports, tests, runbooks: nothing licensed, nothing held back.
No surprises
Fixed price, agreed scope. Overruns are my problem, not yours.
Proof
Three things that were quietly broken.
Anonymised. Details of clients and systems are never disclosed.
Silent data loss under concurrency
A weekly knowledge-base sync fell a full week behind without a single alert: a cloud concurrency cap was silently rejecting one of two ingestion jobs. Diagnosed the race, added bounded idempotent retries, and covered the timeout path with tests.
An unmonitored front door
A production AI service had alarms on a legacy path, and none on the API serving live traffic. Added alarms at both layers and replaced a fourteen-widget dashboard nobody read with a single health view.
Extraction accuracy in document AI
A document pipeline silently dropped line items whenever real-world wording didn’t match reference values. Added normalisation and a human-review fallback: invisible data loss became a visible, reviewable queue.
Who this is for
- You have an AI pilot that demos well but hasn’t shipped
- Your engineers are strong, but new to LLM failure modes
- You’re about to fund an AI build and want independent eyes first
- You’ve shipped one, and aren’t sure you’d know if it broke
- The board keeps asking when the pilot ships, and nobody can answer with evidence
- Not a fit: building a model from scratch
- Not a fit: prompt-engineering workshops
FAQ
Can you help get our AI pilot into production?
That is the entire practice: independent AI consulting for teams whose LLM, RAG or agentic-AI pilot works in demos but hasn’t shipped. The readiness review tells you what stands between the pilot and production; the sprint closes it.
Do you need access to our code?
Read access to code and infrastructure, yes: the findings are evidence-based, not interview-based. NDA first, standard.
What about confidentiality?
NDA before anything is shared. Case studies stay anonymised; client names are never disclosed, referenced or hinted at.
What if the system turns out to be fine?
Then you get a short report saying so, with the evidence. Independent confirmation before a launch is a good outcome, not a wasted one.
Which stacks do you cover?
The major clouds: AWS, Azure and Google Cloud, in Python and TypeScript. The review method doesn’t depend on the vendor. Beyond LLMs: traditional ML, computer vision, OCR, IoT and robotics.
Where are you based, and do time zones work?
Melbourne, Australia; engagements are remote. My mornings overlap US afternoons, my evenings overlap UK and European mornings, and New Zealand is two hours ahead. Calls happen in your business hours, not mine.
Get in touch
Let’s find out what stands between your pilot and production.
Tell me what your system does, what breaks, and what happens when it breaks. If a review isn’t the right thing, I’ll say so.