The Harden-Scale-Operationalise Playbook: Shipping Enterprise AI on Schedule

October 6, 2026
Read Time: 10 min

The model works. The POC is approved. Production is on the roadmap. Six months later, the system is still not live - not because the technology failed, but because nobody had a clear plan for what happens after the demo.

This article is that plan. It covers the three-phase delivery framework Loginsoft uses across enterprise AI deployments, the regulatory requirements that differ by sector, the cost model the POC business case almost always underestimates, and the operating model question that ends more deployments than any technical failure: who owns this on Monday morning?

Every section below draws on the same production engineering practice described in Loginsoft's nine-layer production AI stack - this article is the execution playbook to accompany that architecture reference.

Key Takeaways

  • Productionisation is three distinct phases, each with a defined exit criterion. Treating the move from POC to production as a single sprint is the most common cause of blown timelines. Harden, Scale and Operationalise each have different work, different owners and different done criteria.
  • The real production cost is 3–5x higher than the POC business case estimates. Inference is typically the smallest line. Engineering run-rate, human review capacity, observability tooling and security infrastructure together frequently exceed inference cost at mid-scale deployments.
  • The operating model question must be answered before go-live, not after the first incident. When the system misbehaves at 2am, who gets paged? If that answer does not exist before launch, you do not have an operational system. You have a deployed and unwatched one.

The Three-Phase Delivery Framework: Harden, Scale, Operationalise

The path from AI POC to production is not a single sprint. It is three distinct phases with a defined exit criterion for each. The exit criterion is not a feeling of confidence - it is a documented, tested condition.

Phase Duration Core activities Exit criterion
1. Harden 4–8 weeks Threat modelling, failure-mode mapping, security controls, evaluation harness build, model card, integration design Every known failure mode has a defined, tested behaviour
2. Scale 6–10 weeks Infrastructure sizing, CI/CD pipeline, auto-scaling, model versioning, shadow deployment, load testing System performs within SLA at 95th-percentile production load
3. Operationalise 4–6 weeks (ongoing) Monitoring, alerting, retraining pipeline, runbooks, on-call setup, human-in-the-loop workflows On-call rota exists, runbooks cover known failure modes, first retraining cycle is scheduled
Harden-Scale-Operationalise

Phase 1: Harden - the phase that produces no visible output a board deck can celebrate

Hardening produces artefacts: a threat model, a failure-mode catalogue with defined responses, a model card, an evaluation test suite and a security control specification. The exit criterion is not "we feel confident about security." It is "every known failure mode has a documented, tested behaviour."

Phase 2: Scale - infrastructure at the 95th percentile, not the demo load

Infrastructure design should be benchmarked against the 95th percentile of expected production load. Auto-scaling configuration, model inference concurrency limits and vector index performance should all be tested at that load before shadow deployment begins. CI/CD pipelines for model and prompt updates should be operational before canary begins.

Phase 3: Operationalise - deployed and watched, not deployed and unwatched

Operationalisation is the ongoing work after go-live: the monitoring dashboard, the alert thresholds, the retraining schedule, the human review queue and the on-call runbook. A system without these is not operational.

What slows this down and what you can parallelise

The honest blockers: data access approvals (3–8 weeks at many enterprises), security review scheduling (2–4 weeks), legal review for regulated deployments (variable) and change management for business teams who will own the system. None of these are engineering problems and none of them shorten under timeline pressure. What you can parallelise: evaluation harness build runs concurrently with security control implementation; infrastructure design runs concurrently with integration work.

Loginsoft has run the Harden-Scale-Operationalise programme across enterprise AI deployments in cybersecurity, financial services and engineering operations. We know where the blockers are and build around them from week one.

See how we work

The Bar Moves by Industry: What Regulation Adds to Production-Ready

Sector Regulatory driver Additional release gate conditions
Financial services SR 11-7, EU AI Act (high-risk), MiFID II Model risk management documentation, explainability requirements, independent model validation, adverse action notice capability
Healthcare FDA SaMD guidance, HIPAA Clinical evidence documentation, patient data handling controls, audit trail for clinical recommendations, no-autonomous-action constraints
Manufacturing / industrial ISO 13849, machinery directive Failure mode and effects analysis, safety-rated human override, operational design domain specification
Retail / consumer GDPR, CCPA, EU AI Act (limited risk) Transparency notices, opt-out mechanism, data retention controls, bias audit for personalisation systems

The table above is a starting point, not legal advice. Specific requirements depend on the use case, jurisdiction and risk classification of your system under applicable frameworks.

What Production Actually Costs: The Numbers Missing from the POC Business Case

The POC business case typically contains two cost lines: API credits and engineering time. The production cost model contains seven.

Cost line Driver Order of magnitude
Inference Token volume × model tier $0.002–$0.06 per 1K tokens (input/output blended)
Retrieval / vector infrastructure Query volume, index size $200–$2,000/month at mid-scale
Evaluation runs Harness size × frequency $500–$5,000/month depending on judge model
Observability Trace volume, storage, tooling $300–$3,000/month
Human review Review rate × fully-loaded agent cost 5–15% of queries at launch; highly variable
Security tooling Guardrail evaluation, audit log storage $500–$2,000/month
Engineering run-rate Ongoing ownership 0.5–1.5 FTE depending on complexity
Production AI Cost Breakdown

Cost-per-outcome example: A customer support AI handling 10,000 queries per month at 800 tokens average, using a mid-tier model at $0.01/1K tokens, costs $80/month in inference alone. Add retrieval, observability and a 10% human review rate at $50 per reviewed query, and total cost of ownership reaches ~$5,880/month. At a 70% resolution rate, the system breaks even at 560 resolutions per month - well below the 10,000-query volume. The AI TCO model belongs in the business case before the POC, not after the first quarterly review.

The Operating Model: Who Owns Production AI on Monday Morning

The organisational question that ends more AI deployments than any technical failure: when the system misbehaves, who is accountable?

The five roles a production AI system requires

  • System owner: accountable for the business outcome. Typically a product or operations leader.
  • ML/AI engineer: owns the model layer, evaluation harness, retraining pipeline and model quality.
  • Platform/infrastructure engineer: owns deployment, scaling, observability tooling and SLA.
  • Security engineer: owns the threat model, guardrail layer, audit logging and access controls. On-call for security events.
  • Domain reviewer: a subject-matter expert who reviews a sample of outputs and escalated cases. Part-time but defined and scheduled.

On-call runbooks for AI systems

An on-call rota for a production AI system requires runbooks covering: guardrail fire spikes, latency SLA breaches, model quality degradation identified by monitoring (not by user complaints), integration failures, and provider outages. Each runbook defines the detection signal, the first-responder action, the escalation path and the rollback procedure.

Build, buy or partner: an honest decision test

Build when the AI capability is a core differentiator and you have the engineering depth to own it long-term. Buy when the capability is commodity and the operational overhead exceeds the differentiation value. Partner when you need depth you do not have, the build timeline exceeds your window, or regulatory complexity makes ownership prohibitive. For most first production deployments, the answer is: partner for the platform, build the domain layer, buy the observability tooling. Loginsoft's AI engineering practice is structured for exactly that pattern, with a defined handover plan so your team owns the system at the end of the engagement.

Loginsoft's AI Model Validation team provides independent review of your model quality, security controls and governance documentation - the three sign-off conditions no enterprise AI system should skip.

Explore AI Model Validation

The Pre-Launch Checklist: 22 Conditions, Five Domains

DATA

1. Evaluation set built from representative production inputs, independent of the development team

2. Data freshness SLA defined and instrumented for the retrieval layer

3. Data access authorisation model documented (what the system can access, as whom, under what conditions)

4. Upstream data source failure mode defined and tested

MODEL

5. Model version pinned in deployment configuration

6. Routing and fallback logic defined and tested

7. Prompt registry operational with change history and rollback

8. Evaluation harness runs in CI/CD; gate threshold defined

SECURITY

9. Threat model completed covering prompt injection, data exfiltration and tool misuse

10. Least-privilege access implemented for all non-human actors

11. Audit logging operational and queryable by governance stakeholders

12. Guardrail layer tested against documented adversarial inputs

13. Model card completed and reviewed by security and compliance

OPERATIONS

14. Confidence threshold to human-escalation path defined and tested end-to-end

15. Monitoring dashboard live: latency, cost, quality signals, guardrail fire rate

16. Four critical alerts defined with documented response procedures

17. On-call rota established with runbooks for the top five failure modes

18. Rollback procedure documented and tested before canary begins

COMMERCIAL

19. Cost per outcome calculated at expected production volume

20. Break-even analysis completed and reviewed by the business owner

21. TCO model includes engineering run-rate, not just infrastructure

22. Business owner sign-off on success metrics and review cadence

FAQs

Q1. What does it actually cost to run an AI system in production?

Inference is typically the smallest line in the production cost model. Engineering run-rate (0.5–1.5 FTE for ongoing ownership), human review capacity and observability tooling frequently exceed inference cost at mid-scale deployments. Build a full TCO model before the production business case is approved - not after the first quarterly review.

Q2. Who needs to sign off on an AI system before it ships?

At minimum: the engineering lead (system operates within defined parameters), the security reviewer (threat model addressed, audit logging operational) and the business owner (success criteria defined, escalation paths tested). For regulated deployments in financial services, healthcare or high-risk EU AI Act categories, legal and compliance sign-off is a fourth gate. The sign-off list should be defined before the Harden phase begins.

Q3. Do we need a dedicated AI engineering team, or can existing software engineers handle it?

Existing software engineers can own significant parts of the production AI stack - particularly integration, infrastructure and observability. The gaps typically appear in evaluation methodology for non-deterministic output, LLM-specific security threat modelling (prompt injection, tool misuse) and retraining pipeline design. Most enterprises address this by augmenting a capable software team with specialist AI engineering support during the Harden and Scale phases, then transitioning ownership once the operating model is established.

Q4. How long does the Harden-Scale-Operationalise programme take?

For a well-scoped system with clean data access, 14–24 weeks. The most common elongating factors are data access approvals (3–8 weeks), security review scheduling (2–4 weeks) and change management for business-side owners. Engineering work that can be parallelised - evaluation harness build concurrent with security controls, infrastructure design concurrent with integration work - should be to protect the timeline.

Have a POC that works but hasn't shipped?

Loginsoft delivers a production-readiness assessment across the nine stack layers and four risk dimensions - returned as a gap report with a phased implementation plan. Most assessments complete in two weeks.

Request a Production-Readiness Assessment
Table of Contents

Final Week of September: CISA KEV Grew as Emergency Warnings Dropped Before Advisories

Stay Ahead

Get the Latest Cybersecurity Insights

Security research, threat intelligence, vulnerability updates, product news, and expert insights, delivered directly to your inbox. Stay informed. Stay secure.