Sam Poursoltan

AI consulting · Production reliability

I take AI systems from demo to production.

Most AI pilots stall between proof-of-concept and P&L. I find exactly what will fail, and hand your team the plan to fix it. Fixed fee, five days.

Sam Poursoltan · PhD, Computer Science · 15+ years consulting
Melbourne, Australia · working with teams across AU, US, UK, EU and NZ

Consulting experience with teams at

  • Microsoft
  • AWS
  • IBM
  • NTT
  • Capgemini
  • Insight
  • Rackspace
  • Hewlett Packard

The problem

Your pilot demos well. That was never the hard part.

80.3% of enterprise AI initiatives fail to deliver their intended business value. RAND · 2,400+ initiatives
~95% of generative-AI pilots return nothing measurable on the P&L. MIT · Project NANDA
>50% are abandoned outright after the proof-of-concept stage. Same research

The cause is rarely the model. It is the gap between a demo that works and a system you can trust when nobody is watching it: retries that silently don’t fire, extraction that drops records without an error, an entry point with no alarm on it. Closing that gap is engineering discipline, not AI hype, and it is the shortest path from generative-AI pilot to P&L.

What I bring

The skillset this problem actually needs.

Depth

PhD in Computer Science; Master’s and Bachelor’s in telecommunications and electronic engineering. I read systems to first principles, not to a checklist.

Translation

15+ years consulting, AI lead on medium-to-large programmes. Findings written for engineers, defensible in the boardroom.

Currency

I build and run production agentic-AI systems on the major clouds every day. This isn’t advisory from someone who stopped shipping.

Agentic AI systems

Multi-step LLM workflows and retrieval pipelines that hold up under real traffic.

Document AI at scale

Extraction that treats messy documents as the normal case, and never drops data silently.

Evaluation, not vibes

Deterministic software you unit-test. Non-deterministic systems you evaluate. I design the evals, baselines and regression gates that turn “is it good enough?” into a measured answer.

Observability that works

Alarms on the paths that serve traffic, and a real answer to “would we know if this broke?”

Production readiness

Retries, idempotency, concurrency, load: the discipline that decides whether AI ships.

Before LLMs, there was ML

Traditional ML, image and video processing, OCR, IoT and robotics: the years of shipped systems that tell me which of your problems actually needs an LLM.

Offerings

Start small. Scale only if it earns it.

After a review

Remediation Sprint

2–4 weeks · quoted upfront

I work inside your team to close the review’s highest-risk findings: shipped, tested and monitored, not just recommended.

Ongoing

Fractional AI Reliability

Monthly · scoped to your systems

A standing eye on what you’ve shipped: reviews of new features, incident support, and someone who already knows your stack.

How to think about the cost

  • Pricing is discussed after a conversation about your project, then fixed and quoted upfront. No meter running, no scope creep.
  • A stalled pilot costs a quarter of team salary, and the credibility of the next one.
  • If the review finds nothing serious, that’s the cheapest good news you’ll buy all year.

Independent

No tooling to sell, no partner kickbacks. The findings are the product.

Evidence over opinion

Every finding is reproducible: a trace, a test, or a failing input.

Yours to keep

Reports, tests, runbooks: nothing licensed, nothing held back.

No surprises

Fixed price, agreed scope. Overruns are my problem, not yours.

Proof

Three things that were quietly broken.

Anonymised. Details of clients and systems are never disclosed.

01

Silent data loss under concurrency

A weekly knowledge-base sync fell a full week behind without a single alert: a cloud concurrency cap was silently rejecting one of two ingestion jobs. Diagnosed the race, added bounded idempotent retries, and covered the timeout path with tests.

02

An unmonitored front door

A production AI service had alarms on a legacy path, and none on the API serving live traffic. Added alarms at both layers and replaced a fourteen-widget dashboard nobody read with a single health view.

03

Extraction accuracy in document AI

A document pipeline silently dropped line items whenever real-world wording didn’t match reference values. Added normalisation and a human-review fallback: invisible data loss became a visible, reviewable queue.

Who this is for

  • You have an AI pilot that demos well but hasn’t shipped
  • Your engineers are strong, but new to LLM failure modes
  • You’re about to fund an AI build and want independent eyes first
  • You’ve shipped one, and aren’t sure you’d know if it broke
  • The board keeps asking when the pilot ships, and nobody can answer with evidence
  • Not a fit: building a model from scratch
  • Not a fit: prompt-engineering workshops

FAQ

Can you help get our AI pilot into production?

That is the entire practice: independent AI consulting for teams whose LLM, RAG or agentic-AI pilot works in demos but hasn’t shipped. The readiness review tells you what stands between the pilot and production; the sprint closes it.

Do you need access to our code?

Read access to code and infrastructure, yes: the findings are evidence-based, not interview-based. NDA first, standard.

What about confidentiality?

NDA before anything is shared. Case studies stay anonymised; client names are never disclosed, referenced or hinted at.

What if the system turns out to be fine?

Then you get a short report saying so, with the evidence. Independent confirmation before a launch is a good outcome, not a wasted one.

Which stacks do you cover?

The major clouds: AWS, Azure and Google Cloud, in Python and TypeScript. The review method doesn’t depend on the vendor. Beyond LLMs: traditional ML, computer vision, OCR, IoT and robotics.

Where are you based, and do time zones work?

Melbourne, Australia; engagements are remote. My mornings overlap US afternoons, my evenings overlap UK and European mornings, and New Zealand is two hours ahead. Calls happen in your business hours, not mine.

Get in touch

Let’s find out what stands between your pilot and production.

Tell me what your system does, what breaks, and what happens when it breaks. If a review isn’t the right thing, I’ll say so.

Based in Melbourne, Australia · Engagements are remote

↑