The 14MB Model Behind a $798K Legal Software Business

The 14MB Model Behind a $798K Legal Software Business

A 14-megabyte open model and the ABA's Formal Opinion 512 point at the same gap: deposition review that extracts, cites, and never leaves the laptop.

The 14MB Model and the Documents Nobody Is Allowed to Upload

On August 13, 2026, Cactus Compute released Needle 2: an open 45-million-parameter model that ships as a single 14-megabyte binary and runs a full session in roughly 28 megabytes of RAM. It decodes 500 tokens per second on a Raspberry Pi 5, and 300 to 700 on sub-$200 Android phones. Someone has it running on an ESP32-S3 microcontroller in about 11 megabytes. It's Apache 2.0 licensed and free to download.

The specs are the least interesting thing about it.

Needle 2 isn't a miniature ChatGPT, and Cactus never claimed it was. They built the model for three jobs: calling tools, controlling devices, and pulling structured data out of messy text. The runtime compiles your schema into a grammar, so the model physically cannot emit malformed JSON. Every response carries a learned confidence score you can set thresholds against. On the benchmarks that measure those specific jobs, it trades wins with models five to seventy times its size.

A model that can't write an essay yet reliably converts a paragraph into a validated data structure belongs in a different category of tool altogether. And the business it points at has almost nothing to do with edge AI.

Build the private workstation for the documents people are reluctant, or professionally prohibited, from uploading anywhere. Deposition transcripts, privileged discovery, insurance claim files, medical intake paperwork, inspection reports from sites with no signal.

The software sits on a Mac or a PC. A user drops in documents. It extracts the dates, people, obligations, events, and exhibits, and every extracted fact links back to the exact page, transcript line, or image region it came from. Nothing gets asserted without a receipt. Every claim waits for a human to approve it, and the default processing path never leaves the customer's machine.

That's the heist: AI that never asks the customer to surrender the evidence.

Here's the opportunity:

🎯
The play: A local-first desktop app that turns deposition transcripts into a reviewable evidence table with page-line citations, and never uploads a document.

The money: 2,000 solo licenses at a $399 blended price is $798,000, plus firm plans at $3,000 a year. AirgapAI already sells local-only transcription at $599.

Inside:
• Eight-week MVP scope for Deposition Desk
• Perpetual pricing from $299 to $5,000/year
• Confidence routing that triages every field
• Four moats that survive commodity models

The cost argument is already dead

There's an obvious version of this pitch that sounds smart and is wrong: cloud AI is expensive, run the model locally, keep the margin.

That argument had a shelf life and it expired. OpenAI's cheapest models now run $0.05 to $0.20 per million input tokens, and batch processing and prompt caching cut that by another 50 to 90 percent. Running structured extraction across a 300-page deposition through a cloud API costs less than the coffee the paralegal drinks while reading it.

The cost argument is already dead

Nobody should build local software to dodge a $14 API bill. The reason to build it is that your customer can't put the document in the cloud at all.

In July 2024 the American Bar Association issued Formal Opinion 512, its first formal ethics guidance on generative AI. The commercially important part: lawyers must understand how a tool handles the data they feed it, and they should obtain informed client consent before entering client confidences into a self-learning AI tool. Boilerplate consent buried in an engagement letter doesn't satisfy it.

The consequence is commercial rather than legal. Every cloud AI tool a firm adopts creates a conversation a partner has to have with a client. A tool that processes locally skips the conversation entirely.

That's the wedge: AI for the documents you aren't allowed to paste into AI.

Why the model should be small and the product thick

A litigation paralegal working through a deposition is doing extraction, not analysis. Witness, event date, topic, what got admitted, and where it lives: Smith Dep. 117:4–118:12.

The hard part is converting unstructured testimony into a schema, preserving provenance down to the line number, and making a correction take one keystroke. Composing a graceful paragraph about the witness barely registers.

A 45M-parameter small language model handles a real slice of that work: routing jobs, classifying passages, populating typed fields, deciding what to invoke next. Local OCR reads the scans. Whisper-class models transcribe audio on the machine. A larger quantized model gets called only where synthesis genuinely demands it. SQLite holds the case. The interface does the thing that actually earns money, which is making every output reviewable.

As the model shrinks, the product around it has to get thicker. Plenty of the local AI pitches circulating this year have that backwards, selling the small model as though it were the product.

What happens after a 300-page deposition lands

What happens after a 300-page deposition lands

An experienced litigation paralegal summarizes a transcript at 20 to 25 pages an hour. A 200-page deposition eats eight hours or more. Firms that outsource the work pay $3 to $10 a page, $300 to $2,000 per transcript, and wait three to seven business days for delivery.

Cloud AI is already dismantling that price. Deposition summary services now advertise roughly two cents a page, returned in minutes. If your plan is to sell faster summaries, you're entering a market whose price has already fallen more than 99 percent and hasn't stopped.

Something else hasn't been solved at all. Damien Charlotin's public database of AI hallucination cases tracked roughly 1,500 court decisions worldwide by mid-2026, more than a thousand of them in the United States, where a party relied on AI-fabricated material and a court responded. Penalties climbed from four-figure fines to $15,000 per attorney in a federal appellate case, and about $109,700 in combined sanctions across two lawyers in an Oregon dispute. In February 2026 an Omaha attorney filed an appellate brief containing 63 citations, 57 of them defective and 20 pointing to cases that never existed. He denied using AI and blamed a computer malfunction before admitting it. The Nebraska Supreme Court suspended him that April.

Summarization has become a commodity; verification hasn't. The product that wins is the one that makes a lawyer willing to sign their name underneath the output, which is a very different goal than reading the transcript fastest.

The first product: Deposition Desk

Don't build "PrivateGPT for Lawyers." That's a category description, and categories don't ship. Build one job.

Unlock the Vault.

Join founders who spot opportunities ahead of the crowd. Actionable insights. Zero fluff.

“Intelligent, bold, minus the pretense.”

“Like discovering the cheat codes of the startup world.”

“SH is off-Broadway for founders — weird, sharp, and ahead of the curve.”

Start free, or unlock everything from $35/month.

Already have an account? Sign in.

Similar ideas

New startup opportunities, ideas and insights right in your inbox.