- Productionisation is three distinct phases, each with a defined exit criterion. Treating the move from POC to production as a single sprint is the most common cause of blown timelines. Harden, Scale and Operationalise each have different work, different owners and different done criteria.
- The real production cost is 3–5x higher than the POC business case estimates. Inference is typically the smallest line. Engineering run-rate, human review capacity, observability tooling and security infrastructure together frequently exceed inference cost at mid-scale deployments.
- The operating model question must be answered before go-live, not after the first incident. When the system misbehaves at 2am, who gets paged? If that answer does not exist before launch, you do not have an operational system. You have a deployed and unwatched one.
The model works. The POC is approved. Production is on the roadmap. Six months later, the system is still not live - not because the technology failed, but because nobody had a clear plan for what happens after the demo.
This article is that plan. It covers the three-phase delivery framework Loginsoft uses across enterprise AI deployments, the regulatory requirements that differ by sector, the cost model the POC business case almost always underestimates, and the operating model question that ends more deployments than any technical failure: who owns this on Monday morning?
Every section below draws on the same production engineering practice described in Loginsoft's nine-layer production AI stack - this article is the execution playbook to accompany that architecture reference.
The Three-Phase Delivery Framework: Harden, Scale, Operationalise
The path from AI POC to production is not a single sprint. It is three distinct phases with a defined exit criterion for each. The exit criterion is not a feeling of confidence - it is a documented, tested condition.

Phase 1: Harden - the phase that produces no visible output a board deck can celebrate
Hardening produces artefacts: a threat model, a failure-mode catalogue with defined responses, a model card, an evaluation test suite and a security control specification. The exit criterion is not "we feel confident about security." It is "every known failure mode has a documented, tested behaviour."
Phase 2: Scale - infrastructure at the 95th percentile, not the demo load
Infrastructure design should be benchmarked against the 95th percentile of expected production load. Auto-scaling configuration, model inference concurrency limits and vector index performance should all be tested at that load before shadow deployment begins. CI/CD pipelines for model and prompt updates should be operational before canary begins.
Phase 3: Operationalise - deployed and watched, not deployed and unwatched
Operationalisation is the ongoing work after go-live: the monitoring dashboard, the alert thresholds, the retraining schedule, the human review queue and the on-call runbook. A system without these is not operational.
What slows this down and what you can parallelise
The honest blockers: data access approvals (3–8 weeks at many enterprises), security review scheduling (2–4 weeks), legal review for regulated deployments (variable) and change management for business teams who will own the system. None of these are engineering problems and none of them shorten under timeline pressure. What you can parallelise: evaluation harness build runs concurrently with security control implementation; infrastructure design runs concurrently with integration work.
The Bar Moves by Industry: What Regulation Adds to Production-Ready
The table above is a starting point, not legal advice. Specific requirements depend on the use case, jurisdiction and risk classification of your system under applicable frameworks.
What Production Actually Costs: The Numbers Missing from the POC Business Case
The POC business case typically contains two cost lines: API credits and engineering time. The production cost model contains seven.

Cost-per-outcome example: A customer support AI handling 10,000 queries per month at 800 tokens average, using a mid-tier model at $0.01/1K tokens, costs $80/month in inference alone. Add retrieval, observability and a 10% human review rate at $50 per reviewed query, and total cost of ownership reaches ~$5,880/month. At a 70% resolution rate, the system breaks even at 560 resolutions per month - well below the 10,000-query volume. The AI TCO model belongs in the business case before the POC, not after the first quarterly review.
The Operating Model: Who Owns Production AI on Monday Morning
The organisational question that ends more AI deployments than any technical failure: when the system misbehaves, who is accountable?
The five roles a production AI system requires
- System owner: accountable for the business outcome. Typically a product or operations leader.
- ML/AI engineer: owns the model layer, evaluation harness, retraining pipeline and model quality.
- Platform/infrastructure engineer: owns deployment, scaling, observability tooling and SLA.
- Security engineer: owns the threat model, guardrail layer, audit logging and access controls. On-call for security events.
- Domain reviewer: a subject-matter expert who reviews a sample of outputs and escalated cases. Part-time but defined and scheduled.
On-call runbooks for AI systems
An on-call rota for a production AI system requires runbooks covering: guardrail fire spikes, latency SLA breaches, model quality degradation identified by monitoring (not by user complaints), integration failures, and provider outages. Each runbook defines the detection signal, the first-responder action, the escalation path and the rollback procedure.
Build, buy or partner: an honest decision test
Build when the AI capability is a core differentiator and you have the engineering depth to own it long-term. Buy when the capability is commodity and the operational overhead exceeds the differentiation value. Partner when you need depth you do not have, the build timeline exceeds your window, or regulatory complexity makes ownership prohibitive. For most first production deployments, the answer is: partner for the platform, build the domain layer, buy the observability tooling. Loginsoft's AI engineering practice is structured for exactly that pattern, with a defined handover plan so your team owns the system at the end of the engagement.
The Pre-Launch Checklist: 22 Conditions, Five Domains
DATA
1. Evaluation set built from representative production inputs, independent of the development team
2. Data freshness SLA defined and instrumented for the retrieval layer
3. Data access authorisation model documented (what the system can access, as whom, under what conditions)
4. Upstream data source failure mode defined and tested
MODEL
5. Model version pinned in deployment configuration
6. Routing and fallback logic defined and tested
7. Prompt registry operational with change history and rollback
8. Evaluation harness runs in CI/CD; gate threshold defined
SECURITY
9. Threat model completed covering prompt injection, data exfiltration and tool misuse
10. Least-privilege access implemented for all non-human actors
11. Audit logging operational and queryable by governance stakeholders
12. Guardrail layer tested against documented adversarial inputs
13. Model card completed and reviewed by security and compliance
OPERATIONS
14. Confidence threshold to human-escalation path defined and tested end-to-end
15. Monitoring dashboard live: latency, cost, quality signals, guardrail fire rate
16. Four critical alerts defined with documented response procedures
17. On-call rota established with runbooks for the top five failure modes
18. Rollback procedure documented and tested before canary begins
COMMERCIAL
19. Cost per outcome calculated at expected production volume
20. Break-even analysis completed and reviewed by the business owner
21. TCO model includes engineering run-rate, not just infrastructure
22. Business owner sign-off on success metrics and review cadence
FAQs
Q1. What does it actually cost to run an AI system in production?
Inference is typically the smallest line in the production cost model. Engineering run-rate (0.5–1.5 FTE for ongoing ownership), human review capacity and observability tooling frequently exceed inference cost at mid-scale deployments. Build a full TCO model before the production business case is approved - not after the first quarterly review.
Q2. Who needs to sign off on an AI system before it ships?
At minimum: the engineering lead (system operates within defined parameters), the security reviewer (threat model addressed, audit logging operational) and the business owner (success criteria defined, escalation paths tested). For regulated deployments in financial services, healthcare or high-risk EU AI Act categories, legal and compliance sign-off is a fourth gate. The sign-off list should be defined before the Harden phase begins.
Q3. Do we need a dedicated AI engineering team, or can existing software engineers handle it?
Existing software engineers can own significant parts of the production AI stack - particularly integration, infrastructure and observability. The gaps typically appear in evaluation methodology for non-deterministic output, LLM-specific security threat modelling (prompt injection, tool misuse) and retraining pipeline design. Most enterprises address this by augmenting a capable software team with specialist AI engineering support during the Harden and Scale phases, then transitioning ownership once the operating model is established.
Q4. How long does the Harden-Scale-Operationalise programme take?
For a well-scoped system with clean data access, 14–24 weeks. The most common elongating factors are data access approvals (3–8 weeks), security review scheduling (2–4 weeks) and change management for business-side owners. Engineering work that can be parallelised - evaluation harness build concurrent with security controls, infrastructure design concurrent with integration work - should be to protect the timeline.

Final Week of September: CISA KEV Grew as Emergency Warnings Dropped Before Advisories
Explore the key security, speed, and performance differences between TLS 1.3 and TLS 1.2
Ready to Find and Fix Your Security Weak Points?
LoginSoft's cybersecurity experts help organizations conduct thorough gap analyses, build prioritized remediation roadmaps, and achieve measurable security maturity improvements.
Schedule a Security Assessment
Hari Charan
A MESSAGE FROM OUR TECHNOLOGY LEADER
The NVD enrichment cutback is not a surprise to us - it’s the inflection point we’ve been preparing for. At Loginsoft, we’ve spent years building the research depth and tooling infrastructure to independently enrich vulnerabilities at scale, with the accuracy and context modern security programs require. LOVI is our answer. Our mission is simple: ensure that no CVE relevant to your environment goes unanalyzed, unscored, or unactioned - regardless of what remains in NIST’s queue.
Key Takeaways
- Standard audit tools do not detect intentionally malicious packages. Tools like npm audit and pip-audit focus on known vulnerabilities, not malware behavior. A newly published malicious package with no CVE or advisory can pass these checks undetected.
- The attack surface is trust, not just code. Typosquatting, dependency confusion, and malicious updates exploit package-manager resolution and implicit trust. Developers and CI pipelines may install packages without verifying their provenance or behavior.
- Registry controls are only one layer of defense. Registry scanning, malware detection, and maintainer 2FA reduce risk, but attackers exploit gaps between these controls. Effective protection combines behavioral analysis, dependency pinning, SBOM governance, and build-environment isolation.
Get the Latest Cybersecurity Insights
Security research, threat intelligence, vulnerability updates, product news, and expert insights, delivered directly to your inbox. Stay informed. Stay secure.
BLOGS AND RESOURCES


