IntegrationsBlogBook a free AI assessment
Industry

Prompt Engineering for Lawyers: A Firm Standard, Not a Personal Skill

Prompt engineering for lawyers at firm level: a shared prompt library, a screen for confidential inputs, a citation check and supervision under ABA Opinion 512.

By Mustafa Najoom»Oct 10, 2026»13 min read»prompt engineering for lawyers
Prompt Engineering for Lawyers: A Firm Standard, Not a Personal Skill

Prompt engineering for lawyers is the practice of writing instructions to a generative AI tool so its output stays inside the right law, the right documents and a format a lawyer can check. At firm level it is a standard rather than a personal skill: a shared prompt library with owners and versions, a screen for what may never go into a prompt, and a verification gate on every citation before work leaves the firm. ABA Formal Opinion 512, issued 29 July 2024, puts the supervision of that work on the firm's managerial and supervisory lawyers.

Why does prompt skill not scale across a law firm?

Because it lives in individual chat histories. One associate's tested prompt never reaches the next matter, two lawyers asking the same question get different output, and the partner reviewing a draft cannot see what was asked or what was pasted in. A firm standard fixes the inputs so that review means something.

The training gap makes this worse. For its 2025 Generative AI in Professional Services Report, the Thomson Reuters Institute polled 1,702 people across four professions in January and February 2025. It found that 64 percent had received no generative AI training at work, and that even in law firms and corporate legal departments, the highest-scoring segments, only 40 and 41 percent of respondents said generative AI training was available. Respondents were in eight countries, 42 percent of them in the US. The publisher sells AI tools and only surveyed people already familiar with generative AI, so take the figures as a signal rather than a measurement.

Lists of ChatGPT prompts for lawyers are written for one person, and they seldom record which tool a prompt was tested on, what data it may take or who checks the result. Those are firm questions. Opinion 512 ties competence under Model Rule 1.1 to understanding a tool's capabilities and limitations, and says that whatever review a lawyer chooses, "the lawyer is fully responsible for the work on behalf of the client."

Five parts: the role and task, the jurisdiction and governing law, the source documents the tool may rely on, the output structure a reviewer will check, and an explicit ban on invented authority. Each part narrows what the tool can get wrong and makes errors easier to see.

In a library template, with example wording in italics:

PartWhat it controlsExample wording
Role and taskWho the tool acts as, one task onlyAct as a junior associate preparing a first-draft issue list for a supervising partner.
Jurisdiction and governing lawWhich law applies, what to flagApply [STATE] law and the agreement's governing law clause. Flag anything that turns on another jurisdiction.
Source documentsWhat the tool may rely onUse only the documents below. If the answer is not in them, say you do not know.
Output structureA shape a reviewer can check fastReturn a table: issue, clause number, exact quoted text, risk, open question.
No invented authorityWhat the tool must never doCite nothing that is not in the documents. Support every claim with a quote or mark it unsupported.

The last part carries the most risk. Anthropic's guide to reducing hallucinations recommends the same moves: permit the model to say it does not know, have it extract word-for-word quotes before analyzing a long document, restrict it to the documents provided, and make it find a supporting quote for each claim or withdraw the claim. The guide also says these techniques reduce hallucinations without eliminating them. The prompt is the first control, never the last.

Break large requests into steps, an issue list first and analysis second, so each output is small enough to check.

The five parts of a legal prompt in a firm library template, each with example wording: role and task, who the tool acts as, one task only; jurisdiction and governing law, which law applies and what to flag; source documents, what the tool may rely on, with permission to say it does not know; output structure, a table a reviewer can check fast; and no invented authority, the part that carries the most risk, citing nothing outside the documents and supporting every claim with a quote. These techniques reduce hallucinations without eliminating them, so the prompt is the first control, never the last.

What should never go into a prompt?

Anything the firm has not cleared for that specific tool. Client identities, facts and drafts belong only in a firm-approved tool whose terms the firm has read, privileged material only where counsel directs the use, and protective-order material only if the order allows it. Each library entry names the band of data it may take.

The screen follows Opinion 512, under which Model Rule 1.6 covers all information relating to the representation, whatever its source, which is broader than privilege. Client consent, personal accounts and the policy that governs them are covered in our guide to shadow AI in law firms.

Every library entry can point to one of five bands:

InputPersonal AI accountFirm-approved tool, terms reviewed, training off
Public law, published guidance, blank templatesA policy call, never with client factsAllowed
Client names, facts, drafts, correspondenceNeverAllowed once risk is assessed and consent is settled
Privileged communications and work productNeverOnly at a named lawyer's direction, with a record
Material marked confidential under a protective orderNeverOnly if the order's AI terms are met
Another client's facts inside a templateNeverNever

A protective order can set its own AI terms, as in Morgan v. V2X (D. Colo., 30 March 2026), so check the order before designated material goes into any prompt. Removing names does not always clear a prompt either: a deal's size and timing can identify a client.

What may go into a prompt, five bands of input against a personal AI account and a firm-approved tool with terms reviewed and training off: public law, published guidance and blank templates are a policy call in a personal account, never with client facts, and allowed in an approved tool; client names, facts, drafts and correspondence never go in a personal account and are allowed in an approved tool once risk is assessed and consent is settled; privileged communications and work product never go in a personal account, and in an approved tool only at a named lawyer's direction, with a record; protective order material never goes in a personal account, and in an approved tool only if the order's AI terms are met; another client's facts inside a template, never in either.

For privileged work, the entry also records the directing lawyer and the approved tool. In United States v. Heppner (S.D.N.Y., 17 February 2026), the court suggested that had counsel directed the use, the tool might arguably have acted as a lawyer's agent, without saying that would be enough, and its confidentiality reasoning turned on the platform's privacy terms. Our analysis of the Heppner ruling covers the rest.

How should a firm verify AI citations before anything leaves the building?

The lawyer whose name goes on the work pulls and reads every authority in a primary source, every quotation is matched word for word, and dates and defined terms are checked against the matter file. Confirming that a case exists is not enough. The gate checks that each authority says what the draft claims.

Opinion 512 names the failures seen so far (citations to nonexistent opinions, inaccurate analysis of authority, misleading arguments) and says lawyers must review output, citations included, and correct such errors before filing. A gate makes that review repeatable:

CheckAgainstWho
Each authority exists and is cited correctlyA primary source with a citatorDrafting lawyer
Each authority supports its propositionThe full text, pin cite readDrafting lawyer
Every quotation, date and defined term matchesThe source document and matter fileDrafting lawyer or paralegal
No missing controlling authority or misleading argumentThe supervisor's judgmentSupervising lawyer, before signing
AI disclosure or certificationThe judge's standing orderSigning lawyer, before filing

Our legal AI hub covers how reliable AI research tools have proved in testing and what happened in the fake-citation sanctions cases. The short version: lawyers were sanctioned for not checking, not for using AI, and disclosure rules vary judge by judge.

Opinion 512 lets review be calibrated: a lawyer who tested a tool's contract summaries on a manually reviewed sample and found them accurate need not necessarily review every document. Recorded tests give each library task that basis.

What belongs in a firm prompt library?

One entry per repeatable task, each with a named owner, the tools it is approved for, the data it may receive, the prompt text with a version history, a test record, the checks it requires and a review date. It is a controlled document, not a shared folder of favorite prompts.

FieldWhat it records
Task and ownerOne task, and the lawyer who answers for its output
Approved toolsThe only tools the prompt was tested on
Data bandWhich inputs from the screening table it may take
Prompt text and versionEvery change dated, with the reason
Test recordSample inputs, outputs and the reviewer's verdict
Required checksWhich verification gate rows apply
Ownership Map tierAutomated; Agent-drafted, human-approved; or Human-owned
Retest triggersA new tool, a new model, a failed check, the review date

Build the test record from real past work with a written definition of a passing result, the discipline our guide to evaluating AI agents applies to software. Retest triggers matter because the model under a prompt changes on the vendor's schedule: Anthropic's model deprecations page shows Claude Sonnet 4.5 deprecated on 30 September 2026, with API retirement scheduled for 30 November 2026. Ask each legal AI vendor how it announces model changes.

Templates hold placeholders, never facts: Opinion 512 cites a joint Pennsylvania and Philadelphia bar opinion warning that tools without ethical-wall safeguards may raise conflict issues under Rules 1.7 and 1.9. It also suggests marking AI-generated material in client and firm files, a sensible last step for every entry.

The practice-support lead maintains the library, practice group partners approve their entries, and whoever owns AI at the firm, a steering committee or a Chief AI Officer, approves tools. Our AI steering committee guide covers that structure.

Who is accountable when a library prompt gets it wrong?

The lawyer who relies on the output, every time, whatever review they chose. Opinion 512 also requires managerial lawyers to set clear policies on permitted use, and requires supervisors to make reasonable efforts, including training, so that lawyers and nonlawyers comply.

The library turns each supervision duty into a control it owns:

  • Policy. The firm's AI policy points to the library: library prompts first, approved tools only.
  • Training. Lawyers and staff train on the library's own entries, including each tool's limits and the data band each entry allows.
  • Vendor diligence. An entry lists only tools that passed the firm's vendor file.
  • Client communication. Entries flag tasks whose output feeds a significant client decision, so the responsible lawyer knows when to consult the client.
  • Billing. Each entry carries a time code for prompting and review, so bills show actual time.

The rules behind those controls, from vendor contracts to telling clients and billing AI-assisted work, are answered in our legal AI hub. The Model Rules bind only as each state adopts them, and ethics opinions are advisory, so check your own state's rules. This is information, not legal advice.

When should a prompt become a supervised agent?

When the same library prompt runs on the same kind of document many times a week, needs files from your document management system, and has a verification step that eats most of the time it saves. At that point access control, logging and testing belong in software the firm owns, not in each lawyer's chat window.

Other signals: the pasting is where screening fails, part of the gate can be structured (every citation extracted into a worksheet the lawyer checks), and a client audit may ask who approved what. Then apply the Rent-vs-Own test. Rent a point tool for narrow, common work, such as research inside a vendor's own content. Own a supervised agent when the workflow touches your systems of record, client data or risk, such as your documents and playbooks.

Reference build: illustrative, not a client engagement. A citation-check agent reads a draft brief or research memo from the document management system, only for users who already have access to the matter. It extracts every authority, quotation, date and defined term into the verification-gate worksheet, then fills each row with the passage it found in the matter file or the firm's licensed sources, and marks any row it cannot support. On the Gaper Ownership Map, extraction is Automated, the worksheet is Agent-drafted, human-approved, and reading each authority and signing the filing stay Human-owned. Prompts, tests and approvals stay in the firm's own repository and logs.

Gaper is the AI-native implementation partner that deploys supervised AI agents you own. Our forward-deployed engineers build inside your systems, repository and cloud, against your data and access controls. The Gaper method runs Assess, Scope, Build, Supervise, Hand over, and the firm keeps the code, prompts, evaluation set, runbook and audit trail.

If a well-run library inside a tool you already license covers the work, we will say so; our page on what an AI-native implementation partner does explains when a partner is the wrong call. When one workflow has outgrown its prompt, a free AI assessment maps it and makes the build or buy decision.

Where should a practice-support lead start?

With the prompts people already use. Collect them without penalty, pick the handful of tasks that recur most, write each one up as a library entry with an owner and a test, and train everyone against those entries. A short library that everyone uses beats a long one that nobody opens.

Good first candidates are deposition summaries, first-pass clause review and research memo outlines. A prompt with an owner, a test and a gate is a firm asset; one that lives in a single lawyer's chat history is a risk the firm cannot see.

Book a free AI assessment

Thirty minutes, no commitment. We map one workflow, make the build or buy call, and scope the smallest thing worth shipping.

Frequently asked questions

Why does prompt skill not scale across a law firm?
Because it lives in individual chat histories. One associate's tested prompt never reaches the next matter, two lawyers asking the same question get different output, and the partner reviewing a draft cannot see what was asked or what was pasted in. A firm standard fixes the inputs so that review means something.
What goes into a good legal prompt?
Five parts: the role and task, the jurisdiction and governing law, the source documents the tool may rely on, the output structure a reviewer will check, and an explicit ban on invented authority. Each part narrows what the tool can get wrong and makes errors easier to see.
What should never go into a prompt?
Anything the firm has not cleared for that specific tool. Client identities, facts and drafts belong only in a firm-approved tool whose terms the firm has read, privileged material only where counsel directs the use, and protective-order material only if the order allows it. Each library entry names the band of data it may take.
How should a firm verify AI citations before anything leaves the building?
The lawyer whose name goes on the work pulls and reads every authority in a primary source, every quotation is matched word for word, and dates and defined terms are checked against the matter file. Confirming that a case exists is not enough. The gate checks that each authority says what the draft claims.
What belongs in a firm prompt library?
One entry per repeatable task, each with a named owner, the tools it is approved for, the data it may receive, the prompt text with a version history, a test record, the checks it requires and a review date. It is a controlled document, not a shared folder of favorite prompts.
Who is accountable when a library prompt gets it wrong?
The lawyer who relies on the output, every time, whatever review they chose. Opinion 512 also requires managerial lawyers to set clear policies on permitted use, and requires supervisors to make reasonable efforts, including training, so that lawyers and nonlawyers comply.
MN
Written by

Mustafa Najoom

Marketing & GTM, Gaper

Mustafa is a CPA turned B2B marketer focused on go-to-market strategy, working on growth at Gaper, the AI-native partner that builds and deploys production AI agents.

Ready to turn AI into execution?

Book a free assessment of one workflow. We map it, make an honest build versus buy call before any code, and if an off the shelf product covers the job we will tell you so.