TL;DR
Six attorneys, twelve paralegals, years of scanned documents nobody could search, and every “where is this?” or “how do we handle this?” landing on the same few people. The firm was capped at about $100K a month because more cases meant more hires. We built an in-house AI assistant loaded with every document and procedure the firm had, with OCR so scanned files became searchable. The firm got back about 50 hours a week, took on more cases with the same team, and went from $100K to over $130K a month within six months.
Context and users
- Paralegals were the real users. They fielded questions all day and worked late to do their own jobs afterward.
- Attorneys needed documents and procedural answers fast and didn’t care where they came from.
- The partners wanted growth without another round of hiring, training, and desk space.
The problem, and how we found it
The partners asked for “AI to help with documents.” The assessment showed the document problem was really two problems stacked on each other:
- Scans aren’t searchable. A scanned page is a picture. Years of case files existed only as pictures, so finding anything meant a person who remembered where it was.
- Procedures lived in people’s heads. How the firm handles a given filing was tribal knowledge. New staff learned by interrupting experienced staff.
Both problems had the same symptom, everyone asking the paralegals, and the same cost: overtime, mistakes, and a hard cap on caseload.
Hypothesis
If every document and every procedure lived in one assistant that anyone at the firm could ask, and scanned files were made searchable, the paralegals would stop being the firm’s search engine, the hours would come back, and the firm could take more cases without hiring.
Scope, specs, and prioritization
The order was deliberate: make the archive searchable first, because nothing else worked until the scans were readable; then the procedure answers; then drafting.
| Shipped, in order | Deferred |
|---|---|
| 1. OCR the scanned archive so every page is searchable | Client-facing intake chatbot |
| 2. “Where is this?” and “how do we handle this?” answers from the firm’s own files | Automated filing or e-signature |
| 3. First drafts of letters and forms for an attorney to review | Anything that sends output to a client without attorney review |
Specs that shaped the build:
- Answers cite their source. Every answer points at the document it came from, so a paralegal can check it in one click. This was the accuracy safeguard and the trust builder.
- Attorney review on drafts. The assistant drafts; a lawyer signs. Non-negotiable in a law firm.
- Lives inside the firm’s systems. No new place to log in.
Key decisions and tradeoffs
- OCR first, even though it was the least visible work. Nobody gets excited about text extraction, but shipping the assistant before the archive was searchable would have produced an assistant that couldn’t find half the firm’s files.
- Citations over conversational polish. A slightly clunkier answer with a link to the source beat a smooth answer nobody could verify.
- Drafting last. Drafting was the feature the partners were most excited about, and the one with the most risk. Sequencing it third meant the firm trusted the assistant’s search and answers before it trusted its writing.
What shipped
Andres built the assistant and the OCR pipeline over the firm’s document archive and procedure library, inside the firm’s own systems. I ran the discovery, wrote the scope and the sequencing, defined the success metrics with the partners, and ran the weekly check-ins through launch.
Metrics and outcomes
| Before | After (within 6 months) | |
|---|---|---|
| Monthly revenue | ~$100K | $130K+ |
| Hours back per week | 0 | ~50 (more than one full-time person) |
| New hires needed | Growth required hiring | 0 |
| Paralegal overtime | Routine | Dropped |
| Time to find a document | Ask a person, wait | Seconds, with the source cited |
The 50 hours was the mechanism; the $30K a month was the outcome. The partners used the freed capacity to take more cases rather than to cut staff, which was the plan from the assessment.
What I’d do differently
- Measure interruptions, not just hours. “Questions asked of paralegals per day” would have been a sharper before/after than the hours estimate.
- Bring the paralegals into scoping earlier. The partners scoped; the paralegals used. Two of the best procedure questions came from them after launch and could have been in v1.
- Define “done” for the OCR pass up front. Old scans vary wildly in quality. A clear accuracy bar would have saved a round of rework.
My role vs. the team
I owned discovery, the two-problems framing, scope and sequencing, the metrics, and the partner relationship. Andres owned the OCR pipeline, the assistant, and the integration. The firm’s senior paralegal was the de facto product owner on their side, and the reason adoption was fast.