Insight · Applied AI

Three mid-market AI projects that pay back within 90 days

Counterpoint to the essay on cosmetic AI: three AI projects whose return is measurable within 90 days. Semantic search on a mature catalogue, support operations automation, copywriting under hard constraints. Conditions, metrics, method.

Published June 4, 2026 · 10 min read · Sergio Nokam


In an earlier essay I set out the five artificial intelligence features I refused to build over six months, because the foundations they should have rested on were not in place. The position taken was not anti-AI; it aimed to separate deployments that produce measurable operational value from those that hide structural problems under a cosmetic layer. This essay traces the logical counterpoint: the artificial intelligence projects that, in the mid-market configurations I encounter, repay their investment within a 90-day window.

Three projects come out of the analysis consistently, provided strict eligibility prerequisites are respected. For each, I document below the conditions for success, the observable success metrics, the order of magnitude of the expected return, and the operational method of implementation.

1. Semantic search on a mature catalogue

Strict eligibility conditions. The project requires a catalogue of at least 3,000 genuinely active SKUs, structured product descriptions fully populated on at least 80 % of the volume, a coherent category hierarchy, and enough internal search volume to produce usable usage signal — typically more than 100,000 monthly queries. Absent any one of these prerequisites, the project falls into the category where AI amplifies weak data, and must be deferred.

The operational problem solved. On a mature catalogue, exact keyword search — whether through Elasticsearch, Algolia or native Shopify functionality — suffers a structural flaw: it does not capture queries expressed in the user’s vocabulary rather than the catalogue’s internal nomenclature. A search for comfortable summer t-shirt will not surface items catalogued as lightweight breathable jersey, even though they match the expressed intent. The zero-results rate commonly reaches 20 to 35 % of volume on the mid-market DTC brands I audit.

The calibrated technical solution. Computing vector embeddings on enriched product descriptions (title, long description, attributes, category, brand), via an embedding model suited to French and English — typically a multilingual OpenAI model or a comparable open-source model such as Cohere Embed or Mistral Embed1. Storage in a dedicated vector database — pgvector for brands on PostgreSQL, Pinecone or Weaviate for isolated architectures. Hybrid reformulation of the user query combining classic full-text search with cosine-similarity search in the vector space, weighted empirically.

Observable metrics and expected success thresholds. Reduction of the zero-results rate by a factor of two to four within 60 days of going live. Increase in search-result click-through rate of the order of 15 to 30 %. Measurable conversion lift on sessions that pass through search, generally between 8 and 18 %.

Expected return. For a brand generating $2M in annual revenue attributed to journeys through internal search, a 12 % conversion lift on that segment represents roughly $240,000 of additional annual revenue, or $60,000 per quarter. On an initial investment of the order of $12,500 (the AI Spark offer), the 90-day return is, in nearly every case that respects the prerequisites, better than threefold.

2. Automating second-line support operations

Strict eligibility conditions. The organisation handles more than 500 monthly support tickets, of which at least 40 % concern four to six documentable recurring categories: order status, product return, address change, warranty support, lost delivery, and equivalent cases. The support team has a documented internal policy reference, or can formalise one within the project. The underlying operational systems — order management, customer CRM, 3PL, carrier — expose usable APIs.

The operational problem solved. Second-line support — that is, questions that are not trivial repeated requests but that require reading the customer context and consulting several systems to formulate an answer — typically consumes between 40 and 60 % of the time of mid-market brand support teams. A significant share of these interactions nevertheless follows reproducible patterns amenable to automation, provided that automation respects the brand voice and the commercial sensitivity of the context.

The calibrated technical solution. An architecture of specialised agents on four to six strictly bounded use cases, each with an explicit policy — if situation X, offer A or B, otherwise escalate. Direct integration with operational systems through function calling to retrieve customer context in real time (order status, history, applicable return rights, and so on). Explicit guardrails on sensitive subjects: any case involving a refund above a defined threshold, any marked dissatisfaction, any mention of a legal dispute is automatically escalated to a human operator, with no option for autonomous agent decision.

Observable metrics. Self-resolution rate on targeted cases — the proportion of tickets in covered categories actually resolved without human intervention — between 50 and 70 %. Customer satisfaction (CSAT) on tickets handled by the agent: maintained at or above the human baseline. Average resolution time on covered cases: from the order of 6 to 24 hours down to under 5 minutes.

Expected return. For a support team of four full-time operators each costing $70,000 annually fully loaded (salary + tools + management), a 35 % reduction in volume handled represents the equivalent of 1.4 FTE freed, or roughly $98,000 of annual capacity redeployable to higher-value activity — retention, VIP service, process improvement. On an initial AI Spark investment of $12,500, the 90-day return comfortably exceeds the initial cost, and the freed capacity compounds across later years.

3. Copywriting assisted under hard constraints

Strict eligibility conditions. The brand has a formalised copy brief defining the editorial voice on at least ten concrete dimensions: register, permitted vocabulary, forbidden vocabulary, treatment of technical arguments, standard structure, closing signature. A sample of 50 product records has been written manually against that brief, as a reference corpus for calibration and validation. An internal editorial team or an external copywriter remains engaged for systematic post-editing of assisted output.

The operational problem solved. For a DTC brand whose catalogue exceeds 1,000 SKUs, producing product descriptions manually in the required editorial voice represents a substantial marginal cost — between $35 and $65 per record depending on the writer’s expertise, or $35,000 to $65,000 for a catalogue of 1,000 items. Unconstrained generation by a language model is, as set out previously, a poor solution. The middle path — human copywriting for the records of highest commercial value, assisted and constrained generation for the remainder — produces a defensible compromise between cost and quality.

The calibrated technical solution. A four-stage pipeline. Stage one: structured extraction of product attributes from the PIM or database. Stage two: draft generation by a language model (typically Claude or GPT-4) with a constrained prompt that embeds the brand brief, manual examples from the reference corpus, explicit bounded-length constraints, forbidden vocabulary, and mandatory attributes to mention. Stage three: automatic rule-based validation — length check, forbidden-vocabulary detection, textual similarity check against neighbouring records to identify duplicate-content risk. Stage four: systematic human post-editing before publication, with an average measured time between 5 and 12 minutes per record, against 40 to 60 minutes for writing from scratch.

Observable metrics. Reduction of average time per record by 60 to 75 % against pure manual writing. Maintaining inter-record textual similarity below 40 %, the threshold under which Google does not penalise for duplicate content. Qualitative validation by sampling: across 100 generated records, the copy editor must rate the result acceptable without major rewriting in at least 75 % of cases after prompt calibration.

Expected return. For a catalogue of 1,000 records to produce or rework, the time saved represents the equivalent of $30,000 to $45,000 in writing fees. On an initial AI Spark investment of $12,500 (including prompt calibration and pipeline setup), the return is immediate across the first few hundred records produced.

The AI projects that pay are the ones that graft onto existing foundations — structured data, a formalised editorial voice, documented operations. Not the ones that claim to replace them.

The common prerequisite, and the guarantee that follows

The three projects set out above share one eligibility condition: each requires a pre-existing foundational asset — a mature, structured catalogue for the first, documented and instrumented operations for the second, a formalised editorial voice for the third. Without that asset, none of the three delivers the expected return, regardless of the quality of technical execution.

That condition explains the pre-qualification criterion I apply systematically to AI Spark engagements: before signing, I examine whether the brand actually holds the amplifiable asset. If it does, the engagement is accepted with the contractual AI Spark guarantee — money back or rework at no additional cost if the target KPI is not reached within 30 days of going live. If it does not, I propose upstream work to establish the missing asset — typically the scope of a Diagnostic Stack or an AI Quickstart — before considering any AI deployment.

That ascetic discipline of pre-qualification is the condition of the guarantee. It explains why I can commit to measurable results where other providers, who accept engagements indiscriminately, must dilute their commitments into general execution clauses.

Mid-market AI in 2026 is profitable when it fits that discipline. It is not otherwise. The distinction has become, to my mind, the main dividing line between deployments that create durable value and those that turn out, on retrospective examination, to have been technological representation spending.


Footnotes

  1. On multilingual embedding models suited to French-English: OpenAI text-embedding-3-large, Cohere embed-multilingual-v3, Mistral mistral-embed, BAAI bge-m3. See the public benchmarks of the Massive Text Embedding Benchmark (MTEB): huggingface.co/spaces/mteb/leaderboard.

Frequently asked questions

Which artificial intelligence projects pay back within 90 days?

Three come up consistently in a mid-market context: semantic search on a mature catalogue, automation of second-line support, and copywriting assisted under hard constraints. They share one eligibility condition, which explains their profitability: each grafts onto a pre-existing asset it amplifies — a structured catalogue, documented and instrumented operations, a formalised editorial voice. Without that asset, none of the three delivers the expected return, however good the technical execution.

At what catalogue size does semantic search become profitable?

The project requires at least 3,000 genuinely active SKUs, structured descriptions populated on at least 80 % of the volume, a coherent category hierarchy, and more than 100,000 monthly internal queries to produce usable usage signal. Below those thresholds the deployment amplifies weak data and should be deferred. When the prerequisites are met, the zero-results rate — commonly 20 to 35 % on mid-market DTC brands — drops by a factor of two to four within 60 days, with a conversion lift of 8 to 18 % on sessions that pass through search.

What share of support tickets can be automated without degrading customer satisfaction?

On the categories actually targeted, observed self-resolution runs between 50 and 70 %, provided the organisation handles more than 500 monthly tickets of which at least 40 % fall into four to six documentable recurring categories. Average resolution time drops from six to twenty-four hours to under five minutes, and CSAT holds at the human baseline. The non-negotiable condition is the guardrails: any refund above a defined threshold, any marked dissatisfaction and any mention of a dispute is escalated to a human operator, with no autonomous decision by the agent.