Before you spend on legal AI, build the system
Before a firm spends millions on legal AI, it needs a rule for what to buy, what to pilot, and what should remain human.
Kirkland & Ellis has committed a reported $500 million to proprietary technology and roughly 180 engineers and data scientists. A 20-lawyer firm cannot match that investment. It should not try.
But before either firm spends its next dollar on legal AI, it has to answer the same five questions:
- The work: What are we trying to move?
- The value: What is it worth?
- The risk: What happens when the system is wrong?
- The review: Who catches the error?
- The owner: Who owns the result after the demonstration ends?
If the large firm cannot answer, it can waste millions building infrastructure around an undefined problem. If the smaller firm cannot answer, it can turn a convenient chatbot into an unmanaged part of client work. Budget changes the scale of the mistake, not the question underneath it.
That is the useful conclusion of this series. Part 1 found the money betting that AI will move inside legal work while practitioners still describe it beside the work. Part 2 found verification as the top concern raised by practitioners: how often is AI wrong, and who catches the error before it leaves the firm? This essay finds the system that connects the two. Large firms are hiring teams to build it. Mid-sized and small firms can assemble a minimum version with a managed workspace, a configured agent, one accountable owner, and one rule for deciding which work is valuable and safe enough to delegate.
The scale changes. The operating discipline does not.
Similar capabilities, different provision
Legal AI vendors have settled on a larger promise than an assistant or an agent. They now sell an operating system. Harvey, valued at $11 billion and used by more than 100,000 lawyers according to its own figures, calls its product "the system through which legal work gets done."
The phrase is useful, but only if the parts stay distinct:
- A tool is the model or application.
- The technical system is the managed workspace, integrations, permissions, data, and logging around it.
- The operating model determines which tools are allowed, what data can enter, where a person checks the output, who owns the workflow, how performance is measured, and who did what and when.
- The legal-AI operating system is the technical system and operating model working together.
By that definition, the tool is widely available and the operating system is not.
Koobo coded a balanced sample of 600 AI comments from the same five practitioner venues used in Part 2 to identify where the AI in each account came from.
The headline is not that the big-law venue describes using AI more. What splits is the source.
Comments in the big-law venue more often describe AI selected and supplied by the firm: enterprise licenses, internal models, or specialized tools wired into document systems. In the four comparison venues, AI more often arrives through a personal chatbot or a feature switched on inside software the firm already uses. The general-purpose chatbot appears in both channels.
Access is not the divide. The system around the access is.
What millions buy, and what they do not
The clearest evidence is in who the largest firms hire. Across 2025 and 2026 they opened a wave of senior AI roles, and the postings describe three different jobs.
Some are pure engineering. Kirkland and Pillsbury advertise for people to design what Pillsbury calls "the operational framework," build retrieval pipelines, and run model infrastructure, with no law degree in the requirements. Some are pure law. Latham & Watkins hired counsel to write the firm's AI governance framework, maintain an inventory of what it runs, and answer the AI clauses clients now write into engagement terms. Some sit above both as strategy: Gibson Dunn runs an internal AI Execution Committee, and Ropes & Gray created a Chief of AI Strategy.
Bloomberg Law counted 16 firms recruiting for these roles at once in mid-2026, paying $200,000 to $440,000 and bidding for the same people as Anthropic, Google, and OpenAI.
Collectively, the postings show the two halves of a legal-AI operating system being assembled together: engineers build the technical system; counsel and strategy leaders establish the operating model that governs it. Kirkland describes the aim as one system that carries a matter from scoping through execution, rather than a scattered collection of point tools. A&O Shearman took another route, staffing a 30-person innovation group and, on Harvey's account, putting the product in front of about 2,000 lawyers a day.
This is not evidence that spending creates adoption. It is a warning that spending without the operating model is premature.
On Harvey's own 2026 benchmark, the best models complete fewer than 10% of legal tasks end to end. The firm's answer to that limit is not simply a better model. It is the profession's existing review chain, from junior lawyer to partner, designed into the workflow. Harvey's adoption guide even tracks a "non-adopter conversion rate," a quiet admission that buying the tool and using it are separate achievements.
Before a large firm approves another platform, integration, or proprietary build, it should be able to show where the proposed work sits on a risk-benefit map, what baseline it will improve, where human review enters, and who will own the workflow in production. If it cannot, the expenditure is not an AI strategy. It is an expensive experiment without a stopping rule.
Smaller firms do not need the same system
The gap outside big law is not a lack of interest. Clio's 2024 survey found 40% of solo firms planning to adopt AI in the year ahead, against 24% of large ones. The pattern is consistent with a gap in infrastructure: adoption is wide and shallow at once.
Clio's 2025 report found 72% of solo practitioners using AI in some form but only 8% using it widely or universally, against 35% of large firms. The independent ABA survey reads the same divide from the other end: AI in office practice reaches 18% of solos and 46% of the largest firms, part of a jump from 11% of all firms in 2023 to 30% in 2024.
In the ABA data, 62% to 64% of solo and small firms name ChatGPT. On Clio's 2026 numbers, 57% of solos and 55% of small firms have no written AI policy. The smaller firm often receives AI as a consumer chatbot or as a feature inside the practice-management, research, or productivity software it already pays for. Koobo's comment record sees those embedded features far more often outside big law than inside it.
That packaged path is real, and it is probably the right default. A smaller firm should buy commodity capability rather than recreate model infrastructure. But an embedded feature or chatbot delivers only the tool. Neither decides whether client data belongs there, requires source checking, places a lawyer at the approval point, or measures whether the time saved survives review.
The answer is not to copy Kirkland on a smaller budget. It is to build the smallest operating system the firm can actually run.
For a mid-sized firm, that might be a managed business AI workspace, two or three approved workflows, a part-time owner in legal operations or knowledge management, and a quarterly review. For a small firm, it might be one approved workspace, one configured agent, one named lawyer, one use-case register, and one mandatory checkpoint before anything reaches a client, counterparty, or court.
For a solo, the named owner and the reviewer may be the same person. That makes the system simpler, not the risk lower. Work whose errors are hard to detect may need an outside review path or may remain assistive rather than delegated.
Start with the work, not the product
The first decision is not which vendor to select. It is whether a defined body of work is worth entrusting to AI at all. Name the work first, then judge the value of moving it and the risk if the system gets it wrong.
Benefit means the operational value of moving that work: its volume, time saved, quality or consistency gained, capacity released, and client value. Risk combines the likelihood of an error with its consequences, the sensitivity of the data, whether the error can be detected and reversed, and whether the output reaches a client, counterparty, or court. The decision uses residual risk, what remains after the firm's real control measures are applied, not the protections promised in a demonstration.
Low-risk, high-benefit work is the strongest buy candidate when a credible product already exists. High-benefit work that remains high-risk after existing control measures should not enter live use. It may justify a controlled test environment: no live client matter, no unredacted client data, narrow permissions, a defined test, and mandatory lawyer review. The purpose is to learn whether stronger safeguards can reduce the risk below the firm's approval threshold. If they cannot, the work remains human. Low-risk, low-benefit work can be accepted when it arrives inside an existing platform. High-risk, low-benefit work should remain human.
This matrix decides whether and how aggressively to pursue the work. Only then does build versus buy become useful. Standardized work served by a credible product should usually be bought. Work that depends on proprietary knowledge or deep integration may justify configuration. A differentiated, high-volume workflow may justify a selective build, but only when the firm has the data, ownership, and review capacity to maintain it.
An expensive build in the wrong quadrant is still the wrong decision.
An agent can coordinate the process without owning the decision
For smaller firms, an agent can orchestrate and document parts of the process. It does not replace engineering, governance counsel, professional judgment, or control assurance. It is also optional: a structured form, a use-case register, and an accountable owner can implement the same process.
Before either version scores a use case, the firm needs an eligibility gate: an approved environment; client and matter restrictions; access, retention, and confidentiality terms; a validation method; a qualified reviewer; an incident path; and measurable conditions for stopping or expanding the work.
Imagine a Legal AI Use-Case Steward running inside a managed business workspace. When someone proposes using AI for a body of work, the agent asks for the workflow, intended outcome, data involved, likelihood and consequences of error, review point, and expected benefit. It applies the firm's risk-benefit rubric, identifies missing control measures, and first recommends pursue, controlled test, or do not pursue. If the work is approved to proceed, it then recommends buy, configure, or selectively build. It produces a short decision record for a person to approve, adds an approved use to the register, and returns later with adoption, correction, incident, and time data.
ChatGPT Work inside a managed ChatGPT Business or Enterprise workspace is one current way to assemble this pattern; a comparable managed platform can serve the same role. The product is the example, not the principle. The principle is an agent working from approved sources and tools, with narrow permissions and an explicit stop before consequential action.
The agent can intake, score, document, remind, and monitor. It cannot approve itself. It cannot decide that client information is safe to use, expand its own permissions, or send consequential output without a human. A subscription does not remove the obligation to examine retention, connected systems, client terms, and the exact workflow.
That is a minimum legal-AI operating system: a modest technical system joined to an operating model the firm can actually follow. It does not require an AI department, but every control measure has to be real and enforceable.
It will not give a ten-lawyer firm the infrastructure of a global partnership. It can give the firm a repeatable way to decide where AI belongs, which is the part that matters first.
The divide is the ability to reduce risk
The legal AI divide is not simply between firms that have AI and firms that do not. The same general models are available to nearly everyone. The divide is between firms that can reduce the risk of valuable AI-enabled work and firms that have only acquired the tool.
Big law is trying to reduce that risk with engineers, governance counsel, strategy leaders, integrations, and review chains. Mid-sized firms can do it by concentrating on a few workflows and configuring what they buy. Small firms can do it with a managed workspace, a repeatable decision process, and a human who remains accountable. An agent can coordinate the records and reminders without becoming the control or the decision-maker.
The practical buying question is therefore not only which tool. It is which work, at what benefit, with what residual risk, inside whose operating system.
The final Legal Industry Intelligence report will turn that question into a decision system: what large, mid-sized, and small firms should buy, configure, pilot, selectively build, or leave human; the minimum controls for each path; and a 90-day operating model a firm can adapt. It will include the risk-benefit rubric, worked legal workflows, an agent specification, and the decision records that make the system repeatable.
Before that report, a firm can take the first step now. The Legal AI Readiness Assessment scores strategy, tooling, workflow, governance, and talent, and identifies the weakest part of the system.
Do not start with another tool. Start with the work, the risk, the benefit, and the person who will answer for the result.
A note on method
The first-party figures are the comment layer of the same five practitioner subreddits the series has used throughout, aggregate public data, January 2024 through June 2026. Part 1 counted posts and Part 2 counted the comments underneath them; this essay codes a sample of those comments and never sets a comment-layer figure against a post-layer one. Six hundred AI comments, 300 from the big-law venue and 300 pooled from the other four, were classified by model coders for whether they describe using AI and, if so, where the AI came from: provided by the firm, embedded in an existing tool, or a personal chatbot.
The three displayed source categories do not exhaust every account of AI use. Where a comment named a general-purpose chatbot without saying the firm provided it, it was coded as personal use, and some of those may be enterprise accounts the comment does not flag. A 90-comment sample was re-coded blind three separate times, and the codes agreed on the source in 93% of cases. Shares are read within each venue, and a difference is only called where the two ends separate at a 95% confidence interval.
Venue participation does not verify a commenter's employer or make the sample representative of firms by size. The codes are aggregate only and no comment text is quoted.
The rest is public and named where it appears. The hiring picture is drawn from the firms' own job postings and from reporting in Bloomberg Law and Law360; the framing quotes are on the record from the firms and their leaders; adoption rates are from Clio's Legal Trends surveys and the American Bar Association's technology survey, each carrying its own population, named wherever it is cited. Vendor figures, including Harvey's benchmark and published metrics, are the vendor's own and are labeled as such. The risk-benefit matrix and operating-model recommendations are Koobo's interpretation of the evidence, not findings measured in the comment corpus. This is an operations read based on public data, not legal advice.