LawQi

Module 3.3 · Topic 1

Foundation Models and Providers

Bottom Line Up Front: Foundation models are large pre-trained AI architectures that enable rapid customization for downstream tasks. Understanding which foundation models exist, their capabilities (knowledge cutoff,…

1.1 What Foundation Models Are and How They Differ

Foundation models are large neural networks trained on billions of parameters that learn broad patterns from massive datasets. Unlike narrow models (sentiment analysis, entity recognition), they demonstrate capability across tasks they were never explicitly trained on. Key differentiators include knowledge cutoff date (when training stopped), context window (how many pages it can process), and training approach. A model trained through March 2026 cannot answer about later events. A 4,000-token window cannot process 50-page contracts in one pass. These constraints shape whether a model fits your task.

1.2 Major Model Families and Their Characteristics

The foundation model landscape includes five major model families (OpenAI, Anthropic, Google, xAI and Meta), as well as several commercial and open source competitors from around the world, collectively offering both general-purpose and specialized models. Most of these organizations release frequent updates and advances, often at such a cadence that keeping a table like the one below up to date is a challenge. While the categories and framing of this table remains useful, the specific references date mostly to early 2025, so should not be relied upon. (Pro Tip: Ask LawQi to review and update the "Major Model Families" table found in Module 3.3)

Characteristic Speed & Cost Tier Reasoning & Depth Tier Multimodal Tier
High Speed, Low Cost OpenAI GPT-4o mini, Anthropic Claude 3.5 Haiku, Google Gemini 2.0 Flash, Meta Llama 3.2 (1B–8B) — Claude 3.5 Haiku, Gemini 2.0 Flash (image, video input)
Balanced Performance OpenAI GPT-4 Turbo, Anthropic Claude 3.5 Sonnet, Google Gemini 3.0, xAI Grok 2 Claude 3.5 Sonnet (reasoning chains), Gemini 3.0 (multimodal reasoning) All support images, video analysis; Gemini supports document understanding
Frontier (Deep Reasoning) OpenAI GPT-5.1 (frontier tier), Anthropic Claude Opus 4.6, Google Gemini Deep Research Extended reasoning, complex legal analysis, adversarial robustness Full multimodal pipeline (text, image, video, audio)
Open-Source Alternatives Meta Llama 3.1 (70B), Mistral Large, DeepSeek (14B–70B variants) Llama 3.1 strong reasoning; DeepSeek competitive on cost Llama Vision (70B), Mistral multimodal variants

Each provider (OpenAI, Anthropic, Google, xAI, Meta) offers models at different scales. Smaller models cost 10-100x less for routine tasks. Larger models handle complex legal reasoning but cost more. Know which provider fits your workflow. In early 2025 we might have recommended OpenAI for broad use, Anthropic for safety-focused work, Google for multimodal analysis, xAI for reasoning depth, and Meta for open-source control. In mid-2026, few would apply such limited boundaries to the model families and most would generally speak to aligning functionalities and deployment considerations, rather than starting from the model family perspective. 

1.3 Open-Source vs. Proprietary Models

Proprietary models like ChatGPT, Claude, Gemini are vendor-hosted and accessed via user-facing app or developer-managed API. Your data goes to their servers; they provide SLAs and support. Open-source models like Llama, Mistral, DeepSeek run on your infrastructure, providing full privacy, and no vendor lock-in, but plenty of operational overhead. Proprietary models offer convenience and infrastructure support, but create data residency concerns. Open-source offers security but generally requires self-hosting. For sensitive work, open-source may be worth the burden. For routine work, proprietary convenience often outweighs privacy tradeoff.

1.4 How to Evaluate and Compare Models

Running side-by-side tests on your actual work is the only way to truly evaluate models. As a user, this can mean running multiple apps. As a developer or law firm IT department, this can mean managing multiple APIs. Either way, here's a process you might use to compare models:

  1. Design a test task: Select a real legal task your firm regularly performs: contract review, discovery document prioritization, legal research synthesis, or opposing counsel letter analysis. Define success: what output quality, citation accuracy, or reasoning depth would you accept?
  2. Run the same task on multiple models: Feed the identical input to comparable models available through ChatGPT, Claude, Gemini, and your incumbent tool. Track time, cost, and output quality. For legal work, assess citation correctness and reasoning transparency.
  3. Assess along your criteria: Does the model miss nuances you need? Does it hallucinate citations? Is its depth sufficient or shallow? Compare not just accuracy but reasoning chain quality, supporting evidence, and confidence indicators.
  4. Iterate and select: Run additional tests on edge cases or high-stakes scenarios. Only after multiple trials should you commit to a model for production use. Cost and speed matter, but for legal work, accuracy and auditability must come first.