← All practices + Practice 03 · 5 shipped, all in prod
Practice · 03 of 04 · AI as a Service

Agentic mailers · MCP · Vision · Editors · Humanisation

Ambitious Ideas Into Intelligent Products.

AI architecture, custom model integration, and agents deployed — engineered for measurable growth, not just implementation. 5 live products, none of them demos: an author-first editor with a provenance ledger, an autonomous freight-brokerage agent, a persona workspace with real-time CRDT collaboration, a humanisation engine on Stripe, and a dental clinic operator that spans email automation and computer vision.

Avg. New Customers Per Product
10,000+
Across shipped AI products (verbatim from AI-services)
AI API Integrations Shipped
75+
OpenAI, Anthropic, Gemini, OpenRouter, Ollama, MedGemma
AI Products In Production
5
Manuscripts.ai · Quikest · AI Humanize · TTMS Mailer · Dental Auto
Monitored Agent Ops
24/7
Every AI action logged, timestamped, defensible
What this practice does

What is the AI as a Service practice?

It is the practice that ships AI as production software, not as a proof of concept. Every case study is a live product with paying users or operational uptake — a provenance ledger for authors, an autonomous freight-brokerage agent running 24/7 at ~$0.01 per cycle, a humanisation engine with Stripe billing, a persona workspace with real-time CRDT collaboration, and a dental operator that combines email routing with fine-tuned computer vision. The guardrails come first; the model comes second.

5 live products Agents · Vision · NLP Provenance ledger 24/7 monitored ops Gemini · Claude · MedGemma Prod, not demos
"

We don't just implement AI; we engineer growth through intelligence — ensuring seamless integration, maximum automation, and measurable impact across every product we build.

Build Smarter · Scale Faster · Lead With AI

Momentum matters

Every week without AI is a week competitors get compounding advantage.

How we work

Scope → Plan → Deliver → Handover.

The same shape whether the model is a 1M-context Gemini editor or a LoRA-tuned MedGemma vision annotator. The audit trail is the deliverable, not an afterthought.

01 · One page

Scope

A one-page brief — what the AI has to do, and where a human still has the final call. Fee letter within four working days.

02 · Guardrails first

Plan

Discovery interviews with the operators who'll live with the model. Confidence thresholds, escalation gates and the "wrong-order guard" are designed before the model is picked.

03 · Prod, not demo

Deliver

The model, the orchestrator, the review queue and the observability layer — shipped as a live product, not a Jupyter notebook. Cost tracked per request.

04 · Full audit trail

Handover

Every AI action logged, timestamped and defensible for compliance and QA. The provenance ledger is a first-class deliverable — not an afterthought.

Full AI Service Stack · From Scoping Call to Managed Ops

The engineering behind "prod, not demos".

Six ways we help teams put AI into production — from a first prototype to a fully-managed retainer. Each category ships with real tools, real chips, real code.

01

Ship New AI Products

From zero to a working prototype in hours, not months — then straight into production-grade features.

No Website → Website Rapid Prototyping (Weeks) AI-Accelerated Prototyping Fast Feature Development AI Chatbots & Support Agents
02

Modernize & Agent-ify Legacy

Old apps, dead internal tools and monoliths — turned into API-first, agent-ready systems.

App → MCP Server Legacy → Agentic Apps APIfication Headless-ify Apps Monolith Decomposition Data Migration & Cleanup Revive Dead Internal Tools Legacy-to-Cloud + Agent-Readiness Audits
03

MCP & Multi-Agent Infrastructure

Production MCP servers and orchestration layers so any agent can drive your software safely.

MCP Dev, Hosting & Monitoring MCP Versioning Multi-Agent Orchestration Agent Eval & Observability Governance & Guardrails
04

Vertical AI Agents & Automation

Purpose-built agents for the workflows that eat your team's time — with measured, reportable results.

Legal Intake Insurance Claims Clinic Docs E-commerce Support Voice Agents Support Deflection Systems Workflow Automation Natural-Language Analytics Personalization & Recommendation Engines Internal Copilots
05

AI Strategy, Team & Managed Ops

An embedded partner for planning, running and improving your AI stack, on retainer.

AI Readiness / Opportunity Audits Fractional CAIO / AI Strategy for SMBs AI Training & Workshops Fractional Product/Eng Team Migration Sprints Managed Agent Operations Hosting & Maintenance Retainers Quarterly AI Health Check Subscriptions
06

Beyond AI: The Full Stack

One partner for the rest of the product too — design, growth and brand, built to the same bar.

Mobile App Development UI/UX & Design Systems SEO / SEM Brand Identity
Industries we serve with vertical AI agents

Ten verticals, one deployment model.

The vertical changes; the shape doesn't. A single-purpose agent, with real guardrails, a wrong-order guard, a human review queue and a full audit trail — deployed inside the tools your operators already live in.

  • + Vertical agent

    Legal intake

    with conflict checks and matter routing

  • + Vertical agent

    Insurance claims

    triage, FNOL capture and adjuster assist

  • + Vertical agent

    Clinic documentation

    and patient follow-up

  • + Vertical agent

    E-commerce support

    with order lookup and returns automation

  • + Vertical agent

    Fintech onboarding

    KYC and dispute handling

  • + Vertical agent

    Healthcare scheduling

    and prior authorisation

  • + Vertical agent

    Logistics

    shipment tracking and exception handling

  • + Vertical agent

    Education

    admissions, tutoring and student services

  • + Vertical agent

    Real estate

    lead qualification and site-visit booking

  • + Vertical agent

    Professional services

    proposal drafting and client intake

Deliverables · What we typically ship

Six shapes the AI deliverable usually takes.

The models change every quarter; the shape of what we ship doesn't. Guardrails, orchestration, review queues, cost tracking and provenance are the constants — the model is just the engine bolted to them.

  1. 01

    Agentic mailers

    Autonomous email agents on top of Microsoft Graph — classifying, extracting load details, matching orders, scoring carriers, negotiating within guardrails and booking loads with a full audit trail.

  2. 02

    Humanisation engines

    Text-refinement APIs that rewrite AI-generated content to mimic human writing patterns — with mode-aware fallbacks, rate limiting and sentence-level alternatives.

  3. 03

    AI editors with provenance

    Author-first editors that log every AI-accepted edit as before-and-after text — exportable as JSON (for machines) or Markdown (for humans). Ready for the moment a publisher asks "which sentences are yours?"

  4. 04

    Computer-vision diagnostics

    Fine-tuned detection models (YOLO11m) with LLM annotators (MedGemma 4B-IT) — deployed via Ollama on M4 Mac mini or cloud K8s, with clinician-facing rationale panels.

  5. 05

    Persona & discovery AI

    Two-stage pipelines that turn a URL or paragraph into a full customer persona with demographics, psychographics, biography and contextualised profile imagery — with a chat surface to interview the persona.

  6. 06

    Model routing & cost engineering

    Dual-model orchestration — a cheap fast model as default, a premium model only for high-stakes calls — with prompt caching, tool-use and per-request cost tracking to keep 24/7 loops economical.

Tech stack · The models we actually run

Real models, real versions, real production hours.

Every LLM, framework and library below runs a live Creative Mantra product. Model routing is deliberate: Gemini as default for cost, Claude Sonnet only for high-stakes calls, MedGemma fine-tuned locally for vision annotation, YOLO11m for detection. Every request cost-tracked.

LLMs (routing)
  • +Gemini 2.5 Pro (1M context)
  • +Gemini 3 Flash
  • +Claude Sonnet 4.6 (prompt cache)
  • +Claude Haiku 4.5 via OpenRouter
  • +GPT-4o-mini
  • +Kimi K2 Thinking
Vision & Medical AI
  • +YOLO11m (fine-tuned)
  • +MedGemma 4B-IT (LoRA rank 16, α 32)
  • +bf16 + 4-bit NF4
  • +GGUF Q4_K_M (~1.5 GB)
  • +Ollama on M4 Mac mini
Frameworks
  • +Next.js 16 App Router
  • +React 19
  • +TypeScript 5
  • +Tailwind 4
  • +TipTap v2
  • +Y.js CRDT
  • +Hocuspocus
  • +Harper.js (WASM)
Data
  • +Supabase PostgreSQL
  • +Row-Level Security
  • +SQLite (libsql 0.17.3)
  • +FTS5 (bm25 + snippet)
  • +Outbox pattern
  • +Vercel cron drain
Integration
  • +Microsoft Graph webhooks (35+ mailboxes)
  • +Vercel serverless
  • +Clerk auth
  • +Stripe billing
  • +OpenAI SDK v6.42
  • +OpenRouter
Agent surface
  • +SSE streaming
  • +MC-XXXX error taxonomy
  • +Structured JSON outputs
  • +9 Vercel cron jobs
  • +Confidence thresholds (0.70 / 0.50)
  • +Wrong-order guard
File I/O
  • +pdf-parse
  • +mammoth
  • +docx
  • +jspdf
  • +epub-gen-memory
  • +wink-nlp
  • +graphology
  • +marked
  • +ulid
NOT in the stack
  • +No vector database
  • +No MCP servers (yet)
  • +No multi-agent framework
  • +No fine-tuning pipeline (except MedGemma LoRA)
  • +No LangChain
  • +No LangGraph
The AI products behind this practice
Manuscripts.ai — AI product by Creative MantraQuikest.ai — AI product by Creative Mantraaihumanize.tech — AI product by Creative MantraT-TMS Agentic Mailer — AI product by Creative MantraDentalAuto — AI product by Creative Mantra

Two client engagements (Quikest.ai, T-TMS Agentic Mailer for Red Dog Logistics, DentalAuto for Krest One Dental) and two in-house R&D products (Manuscripts.ai, aihumanize.tech). Every product above has paying users or operational uptake — none are shelf demos.

FAQ

Questions about this practice.

  1. 01

    Which AI/ML models does Creative Mantra actually ship in production?

    Gemini 2.5 Pro (Manuscripts.ai — 1M context for entire manuscripts), Gemini 3 Flash (Quikest, T-TMS Agentic Mailer classification), OpenAI GPT-4o-mini (Quikest full synthesis), Claude Sonnet 4.6 (T-TMS Agentic Mailer orchestration), Claude Haiku 4.5 via OpenRouter (Dental Auto email classification), MedGemma 4B-IT with LoRA fine-tuning (Quikest imagery, Dental Auto vision annotation), and YOLO11m fine-tuned for caries detection (Dental Auto). All in live products, not proofs of concept.

  2. 02

    What is the "wrong-order guard" and why is it in every case study?

    It is the safety rail that hard-stops an agent when an explicit order ID lookup fails — the difference between an agent you trust and an agent that books the wrong load once and gets switched off. Every agentic system we ship has one; every confidence threshold is tuned through conversation with the operator, not guessed. It is what "prod, not demos" means in the AI/ML practice.

  3. 03

    How do you keep AI cost economical for a 24/7 loop?

    Model selection matters. Gemini as default and Claude Sonnet only for high-stakes orchestration keeps the loop economical enough to run 24/7 (approximately $7/month per active user on T-TMS Agentic Mailer, ~$0.01 per cycle). Prompt caching, tool-use, and structured JSON outputs cut token spend further. Every request is cost-tracked so the operator sees exactly what the loop costs before it scales.

  4. 04

    Do your AI systems require human review?

    Yes — every AI/ML system we ship routes through a human for the final call where the decision carries real-world consequence. Dental Auto uses dual thresholds (0.70 auto-forward, 0.50 review queue floor). T-TMS Agentic Mailer runs in three modes (Off, Test drafts, Live autonomous) so dispatchers keep full control. Manuscripts.ai logs every AI-accepted edit for authorial review. No decision is fully autonomous by default.

  5. 05

    What stacks and cloud do you deploy on?

    Vercel serverless Next.js with Supabase PostgreSQL for most builds; Ollama on M4 Mac mini or cloud K8s for on-device vision (GGUF Q4_K_M quantisation, ~1.5 GB runtime for MedGemma); Microsoft Graph webhooks for email ingestion; Clerk plus Supabase RLS for auth. React 19, Tailwind 4, TipTap v2, Y.js and Hocuspocus for real-time collaborative surfaces.

  6. 06

    How long does an AI/ML engagement take?

    A focused agent (like the T-TMS Mailer): 12 to 20 weeks to first live run, then a continuous improvement loop as the model learns from operator feedback. A full AI product (like Manuscripts.ai or Quikest): 16 to 28 weeks to early access, then iterative shipping against real users. Vision fine-tunes: 8 to 16 weeks for dataset curation, LoRA training, and quantisation for deployment.

Now Accepting New AI Projects

Ready to Build Your AI Product?

Tell us what you're trying to automate or ship — we'll scope an AI audit and show you exactly where the wins are. Send a paragraph about what you're building to hello@creative-mantra.com. Reply within 24 hours; fee letter within four working days.