Two knowledge collections need two different schemas
Fit the fields to the size of the collection, and borrow deal details from the systems that already hold them
A firm's knowledge usually sits in two collections, and firms name them differently. One is the material someone made on purpose and approved for everyone to reuse, such as templates and training, often called know-how. The other is the product of past matters, kept so that lawyers can find it again, often called final work or work product.
Firms often design one classification schema and apply it to both. The two collections differ in size, though, and size decides how granular the fields should be. Every organisation draws the line between them slightly differently, but in most cases the approved material is high quality and small in volume, and the bank of past work is far larger.
Know-how: a short, general schema
Granular tags only pay off when there is something to sift. Apply a field for every deal characteristic to a collection of five approved share purchase agreements and each tag ends up describing one or two documents. A lawyer can read five documents and decide for themselves.
So approved material gets general fields that apply to every document: practice area, office, jurisdiction, type of knowledge, type of contract and legal topic. Then add fields on the contextual health of each document: when it was last updated, who updated it, and who owns the content and answers for it. Some collections also carry a health warning field, for a document produced in very specific circumstances that readers should be careful with.
Past work: matter type first
A bank of past work is a much bigger pile, and a mixed one. Litigation work product and transactional work product differ in kind, so lawyers navigate them by different characteristics.
Most of the fields here belong to a particular type of matter, so classification works differently depending on what kind of matter the document came from. A handful of fields sit over everything, such as governing law and jurisdiction, and there won't be many of them. The value of the collection comes from the type-specific fields, because they describe the actual nature of the transaction.
Borrow the deal details from the matter system
Those deal characteristics often exist already, outside the knowledge system. Large firms tend to run systems that track the matters they've worked on. Some hold only the basics, like the client's name and industry sector. Others go much deeper and record the specific characteristics of each deal.
Take a share purchase agreement with a locked box mechanism. In the document management system it carries a client matter number. The matter profiling system uses the same number to track the matter itself, and it might record which purchase price adjustment mechanism the deal used, locked box or completion accounts. Enrichment joins the two. It takes the matter record and the document record and fuses them, so the document picks up the matter's characteristics as metadata.
Without enrichment, a lawyer searches for "share purchase agreement" and "locked box" and relies on keyword matches. An AI agent has to do exactly the same. With it, the lawyer clicks a filter, and the agent knows where to look straight away.
Point the search at one place
Useful context sits in several systems: matter profiling, the CRM, a people system. Each one wired into the knowledge search makes the implementation more complex, and if one integration fails, the whole thing fails.
A data warehouse rationalises those connections into a single integration point. Many firms already use one to take in information from their other databases. Inside it, you build a data product: a chosen handful of fields from each source, stored and aggregated centrally. The knowledge search points at the warehouse, and only at those fields. That gives you one integration, only the data you need, and one place to reconcile systems that record the same matter differently.
The two collections are qualitatively different, and they differ in volume. One is a small set of approved starting points, where a lawyer can browse what's there and the fields only need to say roughly what each document is. The other is the evidence of work the firm has done, where nobody will ever read the pile and the fields have to do the sifting. Design one schema for both and one of them ends up badly served.