# RubyLLM 2.0: Agents That Wait, and Everything Else Your Providers Can Do

Everything in RubyLLM 2.0: resumable agents with human approval, video, speech, OCR, reranking, provider tools, citations, a cost ledger, batches, caching, seventeen providers, and the upgrade path from 1.16.

RubyLLM 2.0 is the biggest release since 1.0. It starts with tool approval:

```ruby
class RefundOrder < RubyLLM::Tool
  description "Refunds an order"
  parameter :order_id, description: "ID of the order to refund"
  requires_approval # that's all it takes

  def execute(order_id:, tool_call: nil)
    Refunds.issue(order_id:, idempotency_key: tool_call.id)
  end
end

class SupportAgent < RubyLLM::Agent
  chat_model Chat
  tools RefundOrder
end

chat = SupportAgent.create!
chat.ask "I was charged twice for order 42." # stops before refunding

# later, when the user approves
chat = SupportAgent.find(chat.id)
chat.approve(chat.pending_approvals.first)
chat.complete # issues the refund and answers
```

The pending call is saved in your database, so the approval can come tomorrow, from another process, after two deploys.

That's just one piece. 2.0 also generates video and speech, reads scanned documents, reranks search results, runs provider-hosted tools with typed citations, accounts for every provider attempt (including the ones that failed), and works across seventeen providers.

This post covers all of it, one section per feature. Each section links to the guide with the full details.

- [Tool approval and durable agents](#tool-approval-and-durable-agents)
- [The agentic loop](#the-agentic-loop)
- [New operations: video, OCR, reranking, tokens](#new-operations)
- [Speech and transcription](#speech-and-transcription)
- [Provider tools](#provider-tools)
- [Citations](#citations)
- [Files and attachments](#files-and-attachments)
- [Cost and usage](#cost-and-usage)
- [Fallbacks, cancellation, and errors](#fallbacks-cancellation-and-errors)
- [Prompt caching](#prompt-caching)
- [Batches](#batches)
- [Owning the transcript](#owning-the-transcript)
- [Providers and protocols](#providers-and-protocols)
- [Rails table ownership](#rails-table-ownership)
- [Prompt templates](#prompt-templates)
- [Workflows and instrumentation](#workflows-and-instrumentation)
- [Coding assistant skill](#coding-assistant-skill)
- [Upgrading from 1.16](#upgrading)
- [Get it](#get-it)

## Tool Approval and Durable Agents

Refunds, deletions, and customer emails shouldn't happen just because a model asked. A person should see the exact call and its arguments first. Declare `requires_approval` on the tool, and the loop parks instead of running it:

```ruby
chat = RubyLLM.chat.with_tools(RefundOrder)
chat.ask "Refund order 42."

chat.awaiting_approval? # => true, and nothing has executed
call = chat.pending_approvals.first
call.name               # => "refund_order"
call.arguments          # => {"order_id" => "42"}

chat.approve(call)
chat.complete           # runs the refund, then the model answers
```

`ask` runs the loop until there's nothing left it's allowed to do, then returns. No thread blocks waiting for a human: the chat records the request, and your code decides.

`chat.deny(call)` skips the call. The model gets a tool result that reads `The user denied the refund_order tool call.` and can explain, offer something else, or ask what you'd prefer. Nothing raises. Tools that don't need approval still run in the same round, so if the model asks to look up the order and refund it, the lookup happens and the refund waits.

A decision has three values: `true` (approved), `false` (denied), and `nil` (still waiting). If your app already stores approvals, give the tool a resolver block, `requires_approval { |call| ... }`, that returns one of those three. The resolver is then the source of truth. RubyLLM may call it several times while a call waits, including after a restart, so keep it a read.

While a call is unanswered, `ask` and `ask_later` raise `RubyLLM::PendingToolCallsError`. Providers reject a transcript with a tool call that has no result, so RubyLLM raises before sending it.

Pausing in memory is easy. The harder case is a human who takes twenty minutes to answer while you deploy twice. With `acts_as_chat`, the tool call and its decision are database rows, so the approval can come from anywhere:

```ruby
class ApprovalsController < ApplicationController
  def create
    chat = Chat.find(params[:chat_id])
    chat.approve(params[:tool_call_id]) # or chat.deny(...)
    CompleteJob.perform_later(chat.id)
  end
end

class CompleteJob < ApplicationJob
  def perform(chat_id)
    SupportAgent.find(chat_id).complete # resumes where it stopped
  end
end
```

`SupportAgent.find` reloads the transcript and reapplies the agent's tools and configuration. Load through the agent, not `Chat.find`, when you resume: a bare record has the transcript but no tools, so it doesn't know `RefundOrder` needs approval and has nothing to run once it's approved. Tool calls that already have saved results are skipped. Render `pending_approvals` as cards showing each call's name and arguments. There's no workflow engine and no state machine to serialize: the transcript holds the state.

One rule comes with it: tools that write should be idempotent. A process can die after a tool's side effect succeeds and before its result is saved, and the next job runs the approved call again. That's why the refund passes `tool_call.id` to the payment provider as an idempotency key. For writes in your own database, a unique constraint on the call ID does the same job.

Approval also covers provider-run tools. When OpenAI or Azure calls a tool on a remote MCP server, the provider can pause before the call, and it lands in the same `pending_approvals` list with `call.remote?` set. See [provider tools](#provider-tools).

Thanks to [@jondavidschober](https://github.com/jondavidschober) for the issue that started this ([#503](https://github.com/crmne/ruby_llm/issues/503)).

The [tool approval guide](https://rubyllm.com/tool-execution/#requiring-approval) and [durable agents guide](https://rubyllm.com/durable-agents/) cover the full flow, including ActiveJob Continuations for surviving deploys mid-run.

## The Agentic Loop

`ask` still runs the whole agentic loop for you. When you want control, 2.0 breaks it into verbs:

```ruby
chat = RubyLLM.chat.with_tools(SearchDocs).ask_later("How do I configure webhooks?")
chat.step until chat.complete? || chat.awaiting_approval?
```

`ask_later` stages a message. `generate` makes one model call. `run_tools` runs pending tools. `step` picks the next move. `complete` keeps going until there's an answer or a decision to make, and `ask` is `ask_later` followed by `complete`. `cancel` stops a run from another thread, or, in Rails, from another process through the database.

Each verb decides what to do next by reading the persisted messages, so the loop doesn't need to live in one process. You can run one step per job, each with its own retry boundary. `run_tools` skips calls whose results have been saved, so a process that dies after one result out of three resumes with the remaining two. With these verbs, iteration budgets and handoffs between agents each take a few lines of Ruby.

`halt` and `RubyLLM::Tool::Halt` are gone. In 1.x a tool could end the conversation from its return value. In 2.0 your loop decides when to stop, and `chat.with_tool_options(calls: :one)` limits the model to one tool call per response.

I wrote about this in detail in [The Agentic Loop, Exposed](/rubyllm-2-0-agentic-loop/). The [agentic workflows guide](https://rubyllm.com/agentic-workflows/) has the patterns built on these verbs.

## New Operations

1.x could chat, embed, paint, transcribe, and moderate. 2.0 adds video, speech, OCR, reranking, multimodal embeddings, and tokenization:

```ruby
RubyLLM.animate("A red panda typing Ruby code").save("panda.mp4")
RubyLLM.speak("Welcome to the Ruby study group.").save("welcome.mp3")
RubyLLM.ocr("scanned-contract.pdf").markdown
RubyLLM.rerank("How do I reset my password?", documents, model: "rerank-v3.5")
RubyLLM.embed("A red panda", with: "panda.jpg", model: "gemini-embedding-2")
RubyLLM.tokenize("Hello, Ruby!", model: "grok-4.3").ids
```

None of them needs a chat. Results are typed, and `save` and `to_blob` work on images, speech, video, and downloaded files.

**Video.** `animate` submits a video job, polls until the provider finishes, and returns a `RubyLLM::Video`. Pass a still image with `with:`, as in chats, to animate it, and `extend:` to continue a video you generated. `animate_later` returns a `RubyLLM::VideoJob` right away. Call `refresh` on it from any process; `done?` is true once the job finishes either way, `completed?` only when you got a video, and `wait` runs the polling loop for you. Gemini, Vertex AI, xAI, Azure, Bedrock, OpenRouter, ElevenLabs, and GPUStack all sit behind the same call.

**OCR.** `RubyLLM.ocr` returns markdown per page, plus the images and tables the provider found and its raw page data, without involving a chat model. Mistral Document AI reads PDFs, office documents, and images; Cohere Parse reads one image per request. `pages:` takes zero-based page indexes, because every OCR API has that concept. Table formats differ between providers, so options like `table_format` go in `provider_options:`, in the provider's own words.

**Reranking.** Embeddings are good at finding fifty documents that might answer a question. They're worse at deciding which five actually do. A reranker reads the query against each candidate and sorts them, and each result's `index` maps it back to your records:

```ruby
query_vector = RubyLLM.embed(query).vectors
candidates = Article.nearest_neighbors(:embedding, query_vector, distance: :cosine).limit(50).to_a
rerank = RubyLLM.rerank(query, candidates.map(&:body), model: "rerank-v3.5", top_n: 5) # keeps the best 5

best = rerank.results.map { |result| candidates[result.index] }
```

`model:` is required, because reranker catalogs differ per provider and there's no sensible default. Scores belong to the model that produced them: fine for ordering and for a cutoff you tune on your own data, meaningless to compare across rerankers. Cohere, Bedrock, Vertex AI, Azure, OpenRouter, and GPUStack all answer `rerank`.

**Tokens.** `RubyLLM.tokenize` returns the token IDs for exactly the string you pass, on xAI, Cohere, and GPUStack. Usually you want the size of a whole request, so ask the chat: `chat.count_tokens("Summarize this contract.")` counts history, instructions, function tools, schema, thinking settings, and supported attachments through the provider's own counting endpoint, without adding the question to the conversation. `RubyLLM.count_tokens(text, model:)` does the same for a lone prompt. Anthropic, OpenAI, Gemini, Vertex AI, and Bedrock count requests, and anything unsupported raises `RubyLLM::Error` instead of guessing.

Counting endpoints don't receive provider tools, `provider_options`, compaction, or `before_request` edits, so a chat using those can send something slightly different from what you counted. Neither count predicts the output; `response.tokens` has the actual counts afterwards.

Each one has a guide: [video](https://rubyllm.com/video-generation/), [OCR](https://rubyllm.com/ocr/), [reranking](https://rubyllm.com/rerank/), [embeddings](https://rubyllm.com/embeddings/), and [tokenization](https://rubyllm.com/tokenization/).

## Speech and Transcription

RubyLLM has had transcription since 1.x. 2.0 adds text to speech with `RubyLLM.speak`, which returns a `RubyLLM::Speech` with `model`, `voice`, `format`, `mime_type`, `to_blob`, and `save`. With both, a Rails action can answer a voice message with audio:

```ruby
def create
  audio = params[:audio]
  return head :bad_request unless audio.is_a?(ActionDispatch::Http::UploadedFile)

  question = RubyLLM.transcribe(audio).text
  answer = Current.user.support_chat.ask(question).content

  speech = RubyLLM.speak(answer) # text to speech
  send_data speech.to_blob, type: speech.mime_type, disposition: "inline"
end
```

Check that the upload is actually a file: a bare string parameter would make RubyLLM read a path or fetch a URL you didn't intend. The middle step is a normal chat, so it can use tools, approvals, and history. The three steps are ordinary sequential requests, not a realtime audio stream.

OpenAI, Gemini, ElevenLabs, Deepgram, Mistral, xAI, Azure, Vertex AI, OpenRouter, and GPUStack all answer to `speak`. `model:`, `voice:`, and `format:` are keywords because every provider has them. Voices are provider-specific: OpenAI and Gemini use names, ElevenLabs uses voice IDs from your voice library, and Deepgram bakes the voice into the model name. Everything else, like OpenAI's `instructions` and `speed`, ElevenLabs' `voice_settings`, goes in `provider_options:` in the provider's own words. I pass each provider's options through rather than invent a lowest-common-denominator `emotion:` parameter that would cover only part of what each one supports.

Pass a block and the audio streams as `SpeechChunk`s, while the call still returns the complete `Speech`, so you can play first and save after. Two Gemini caveats: its speech endpoint doesn't stream, and asking it to raises instead of quietly buffering; and it returns raw PCM, so `speech.format` is `"pcm"` and naming the file `.wav` doesn't make it WAV. Run it through ffmpeg when you need a container.

Transcription gained streaming and shared controls for speakers and timestamps:

```ruby
transcription = RubyLLM.transcribe(
  "standup.wav",
  model: "gemini-3.5-transcribe",
  speaker_names: [], # speaker labels
  timestamps: :word # per-word timing
)

transcription.words.each do |word|
  puts "#{word['speaker']} at #{word['start']}s: #{word['word']}"
end
```

Speaker labels now work beyond OpenAI, and OpenAI's diarization model can still match voices to names from short reference clips with `speaker_references:`. Pass a block to stream: `delta?` marks committed text, `partial?` a tentative guess the next chunk replaces, and `segment?` speaker and timing data. Deepgram, ElevenLabs, xAI, and Google's live models stream over WebSockets with the optional `websocket-driver` gem; OpenAI, Mistral, and others stream over plain HTTP. Either way, the input is a recording you already have. `format:` asks for the provider's transcript format, such as OpenAI's `"srt"`.

The [text to speech guide](https://rubyllm.com/text-to-speech/) and [transcription guide](https://rubyllm.com/audio-transcription/) have every model, voice, and format.

## Provider Tools

Providers now run tools on their own infrastructure: web search, code execution, file search, remote MCP servers. With `with_provider_tools`, web search needs no search client or extra key:

```ruby
chat = RubyLLM.chat(model: "claude-sonnet-5").with_provider_tools(:web_search) # runs on Anthropic's side
response = chat.ask "What changed in the latest Ruby release?"

response.citations.each { |citation| puts "#{citation.title}: #{citation.url}" }
```

Your RubyLLM tools run in your process, provider tools on the provider's side, and one chat can use both: `chat.with_tools(Weather).with_provider_tools(:web_search, :code_execution)`. The aliases make provider tools portable. `:web_search` becomes Anthropic's versioned search tool, OpenAI's Responses search tool, Gemini's Google Search grounding, and the equivalent on other providers, so you can switch models without changing the code. The full set is `:web_search`, `:web_fetch` (or `:url_context`), `:x_search`, `:code_execution`, `:file_search`, `:image_generation`, `:apply_patch`, and `:mcp`.

Options use the provider's own vocabulary, like `web_search: { allowed_domains: ["ruby-lang.org"], max_uses: 3 }`, because a lowest-common-denominator filter language would lose what each provider supports. The alias works across providers; the options are specific to one. Providers also ship new tools faster than any library can wrap them, so a positional Hash goes into the payload as is. RubyLLM handles the response blocks the protocol returns, so you need a definition the endpoint accepts, not a new gem release.

With `:mcp`, the provider connects to a remote MCP server, lists its tools, and calls them, and your process never opens a connection to it. On OpenAI and Azure Responses, the provider can stop and ask before each call, and those requests land in the same approval flow as your Ruby tools:

```ruby
chat = RubyLLM.chat(model: "gpt-5.6")
  .with_provider_tools(mcp: {
    name: "docs",
    url: "https://learn.microsoft.com/api/mcp",
    allowed_tools: ["microsoft_docs_search"],
    require_approval: "always" # the provider asks first
  })

chat.ask "Search Microsoft documentation for Azure Blob Storage."
call = chat.pending_approvals.first
call.remote? # => true

chat.approve(call) # or chat.deny(call)
chat.complete
```

Tool activity shows up on the response: `response.server_tool_calls` gives each call's name, input, result, and the provider's raw block. Search results are the same `Citation` objects you get from documents, and images from `:image_generation` or files from `:code_execution` come back as `response.attachments`. Providers often bill tool uses on top of tokens, so `response.tokens.server_tool_use` gives you the per-use counters to price them.

Some providers need their tool blocks replayed on later turns, and Anthropic rejects a conversation that drops them. RubyLLM keeps them on the message and sends them back. In Rails, the `server_tool_calls` and `raw_content` columns keep them across requests, and an agent can declare `provider_tools :web_search, :code_execution` so they come back with the chat.

Availability depends on the provider, model, and protocol. Ask for an alias a provider doesn't have and you get `RubyLLM::UnsupportedServerToolError` before any request, listing the ones it does. A few tools need a non-default protocol: Mistral's hosted search and code execution want `protocol: :conversations`, Gemini's remote MCP wants `protocol: :interactions`, and OpenRouter's hosted shell and MCP want `protocol: :responses`. Not every provider can pause for MCP approval either: Anthropic, Gemini Interactions, and xAI run allowed MCP tools immediately, so choose which tools to allow.

The [provider tools guide](https://rubyllm.com/provider-tools/) has the per-provider tables, file search setup, and MCP connection details.

## Citations

When a model tells your user "The contract can be terminated with 30 days' notice," the user needs to know where that comes from and which page to check:

```ruby
chat = RubyLLM.chat(model: "claude-sonnet-5").with_citations # document citations
response = chat.ask "What are the termination conditions?", with: "contract.pdf"

response.citations.each do |citation|
  puts "p. #{citation.start_page}: #{citation.cited_text}"
end
```

`with_citations` turns on document citations for Anthropic, Cohere, and Claude on Bedrock, and `with_citations(false)` turns them off. Web search citations need no setting: turn on the provider's search tool and they show up, and Perplexity's Sonar models search on every request.

Every provider has its own citation format. Anthropic returns citation blocks. Gemini returns grounding metadata with byte offsets. OpenAI hangs annotations off the response text. Perplexity sends a list of URLs on every chunk. If you build footnotes against each one directly, you write the footnote feature once per provider. RubyLLM turns all of them into `RubyLLM::Citation`, with `url`, `title`, `cited_text` (the quoted passage), `text` (the part of the answer it supports), `start_index` and `end_index`, `source_id`, `source_index`, and `start_page` and `end_page`. Whatever the provider didn't report is `nil`; a web citation often has a URL and nothing else.

Positions are character offsets into `response.content`, ready to slice. Gemini reports byte offsets, which would drift the moment your answer contains an accented letter, so RubyLLM converts them and your footnote lands after "Café" and not in the middle of it.

Citations matter most in RAG, which is also where they usually get lost: your tool returns a JSON blob, and the model paraphrases it with no trail back. Return `RubyLLM::SearchResults` instead:

```ruby
class KnowledgeBase < RubyLLM::Tool
  description "Searches the company knowledge base"
  parameter :query, description: "What to look for"

  def execute(query:)
    docs = Document.search(query)
    return "No matching documentation found." if docs.empty?

    RubyLLM::SearchResults.new( # citable results
      *docs.map { |doc| { title: doc.title, url: doc.url, text: doc.body } }
    )
  end
end
```

Anthropic, Cohere, and Bedrock receive these as their native citable search results, including after a Rails chat is reloaded from the database, and `response.citations.first.url` is the `doc.url` you returned. Other providers get the results as JSON text, which they can read but not formally cite.

Citations arrive on streaming chunks, and the final message collects them all, deduplicated. Some providers only send citations at the end, so build the finished footnotes from the final message. With `acts_as_message`, they're saved as JSON and come back as `Citation` objects. One limit: Anthropic won't combine document citations with `with_schema`, so a single request gets either citations or structured output.

The [citations guide](https://rubyllm.com/citations/) has a footnote renderer you can copy.

## Files and Attachments

If a user asks five questions about the same 40 MB PDF, there's no reason to send those 40 MB five times. Upload it once:

```ruby
manual = RubyLLM.upload("manual.pdf", provider: :anthropic) # upload once

chat = RubyLLM.chat(model: "claude-sonnet-5")
chat.ask "Summarize the manual.", with: manual
chat.ask "What does the warranty cover?"

RubyLLM.chat(model: "claude-sonnet-5").ask "Which page explains the reset button?", with: manual
```

The requests point at the provider's copy, and `with:` takes an uploaded file like a path or a URL. This saves the transfer, but you still pay for the document's tokens on every request. [Prompt caching](#prompt-caching) is what makes re-reading it cheap.

`RubyLLM.upload` takes a path, an IO, or a `RubyLLM::Attachment`, and returns a `RubyLLM::UploadedFile` with its `id`, `provider`, `filename`, `byte_size`, `mime_type`, and `expires_at`. A file ID only works with the provider that issued it, so if you store IDs, store the provider next to them and use `RubyLLM::UploadedFile.find(id, provider:)`. `RubyLLM.download(file_id, provider: :openai)` brings files back with the same `save` and `to_blob` as other media. Twelve providers have file APIs behind this, including S3 for Bedrock and Cloud Storage for Vertex AI. Fewer can reference a stored file from a chat, and download rules are the provider's: Anthropic and OpenRouter, for example, only let you download files their tools created.

Most of the time you won't call `upload` at all. Attach the file as usual, and if it's too big to send inline, RubyLLM uploads it and sends a reference instead. Gemini switches over at 20 MB (7 MB on Vertex AI), Anthropic at 24 MB, OpenAI at 50 MB for PDFs and documents. The upload is remembered on the attachment per provider and per set of credentials, so follow-ups reuse it, a provider switch uploads once there, and an expired file is uploaded again. Your local history keeps the original file. The memo lives in memory, though, so a Rails chat reloaded in another process uploads its stored file again. `config.auto_upload_large_files = false` turns this off.

Tools can return files too, so a chart tool can return the chart:

```ruby
class RevenueChart < RubyLLM::Tool
  description "Renders a revenue chart for a quarter"
  parameter :quarter, description: "Quarter, like 2026-Q3"

  def execute(quarter:)
    path = Charts.revenue(quarter).render_png
    ["Revenue chart for #{quarter}", RubyLLM::Attachment.new(path)] # text plus a file
  end
end
```

Strings become the tool result's text and attachments become its files. A browser tool can return a screenshot for a vision model to look at. Anthropic and Bedrock take tool files inside the tool result, Gemini 3 inside the function response, and OpenAI, whose tool results are text-only, gets them in a user message right after the result. You write the return value once. If a provider can't take that file type at all, you get `RubyLLM::UnsupportedAttachmentError` rather than a model answering about a file it never saw.

The [files guide](https://rubyllm.com/files/) covers expiration, provider limits, and storage setup, and the [attachments guide](https://rubyllm.com/attachments/) covers everything `with:` accepts.

## Cost and Usage

Chat history doesn't match what you pay for. A retry that failed, a fallback model that took over, and a stream the user cancelled halfway all cost tokens, and none of them leave an assistant message behind.

2.0 records usage per provider attempt, because providers bill attempts and one message can take several:

```ruby
chat = RubyLLM.chat.with_fallbacks("claude-sonnet-5")
response = chat.ask "Explain Ruby fibers in one paragraph."

response.tokens.input
response.cost.total
chat.cost.total # every attempt, including the ones that failed
```

`response.tokens` and `response.cost` add up every attempt that produced that response: retries, fallbacks, and the one that succeeded. `chat.tokens` and `chat.cost` add up every attempt the chat ever made, including cancelled ones with no message to show for it. One-shot operations like `embed` and `transcribe` return the same objects.

Providers disagree on whether cached tokens are part of the input count, so RubyLLM normalizes the buckets before pricing. `tokens.input` is ordinary input excluding cache reads and writes, and `output`, `cache_read`, `cache_write`, and `thinking` mean the same thing on every provider. When a provider bills thinking as output, it's already in `tokens.output`, so don't add it twice. `cost` mirrors the buckets and adds `total`.

When usage or pricing is unknown, the total is `nil` rather than zero, so a request you can't price doesn't look free. That covers a model missing from the pricing registry and an attempt that may have been billed but never reported usage: a timeout, a server error, a stream that died early. An attempt that provably wasn't billed, like a refused connection or a 4xx rejection before the model ran, records zero. Provider-reported charges from OpenRouter or xAI are kept in `tokens.reported_cost` and win over estimates. Hosted tool uses show up as counters on `tokens.server_tool_use`, for example `{"web_search_requests" => 2}`; unless the provider reports the charge, `cost.total` covers tokens only.

With `acts_as_chat`, every finished attempt is written to `ruby_llm_usages` before the assistant message callback runs, so a cancelled stream leaves an accounting row even when it leaves no message. Each row stores the operation, provider, model, status (`succeeded`, `failed`, or `cancelled`), the token buckets, and decimal costs. Costs are frozen when the attempt finishes, at the prices RubyLLM knew then, so run `RubyLLM.models.refresh` on a schedule; daily is plenty.

The table is good for reporting with one catch: SQL `SUM` skips `NULL`, so a sum of known costs looks like a complete bill even when it isn't. Count the unpriced attempts next to it:

```sql
SELECT model,
       SUM(total_cost) AS known_cost,
       SUM(CASE WHEN total_cost IS NULL THEN 1 ELSE 0 END) AS unpriced_attempts
FROM ruby_llm_usages
WHERE created_at >= '2026-10-01'
GROUP BY model;
```

Every finished attempt also emits a `usage.ruby_llm` event with `operation`, `provider`, `model`, `status`, `tokens`, and `cost`, for embeddings, images, speech, transcription, moderation, OCR, and reranking as well as chat. Use it to feed metrics, billing, or finance systems.

The [Tokens and Costs guide](https://rubyllm.com/cost-and-usage-tracking/) explains exactly what each total includes, and covers pricing usage yourself with `cost_for`.

## Fallbacks, Cancellation, and Errors

Providers go down, and switching the model and redeploying is a slow way to respond. Declare backup models on the chat:

```ruby
chat = RubyLLM.chat(model: "claude-sonnet-5")
  .with_fallbacks("gpt-5.6", "gemini-3.7-flash") # tried in order if Claude fails

response = chat.ask "Summarize this incident report."
```

If Claude fails with a rate limit, a server error, an overload, a timeout, or a dropped connection, RubyLLM retries the same request on GPT, then Gemini, with the same conversation, tools, schema, and settings. Once that generation is done, the chat goes back to Claude. Authentication and bad-request errors don't fall back, because a wrong API key on provider A is not a reason to quietly bill provider B. `on:` replaces the list, and the defaults are a public constant, `RubyLLM::Fallback::DEFAULT_ERRORS`, so you can add `RubyLLM::ContextLengthExceededError` to hand long conversations to a model with a bigger window. Agents declare the same thing with `fallbacks "gpt-5.6", "gemini-3.7-flash"`.

Fallbacks need credentials for each provider, and each fallback has to support what you're asking for, such as structured output. `before_fallback` and `after_fallback` tell you when it happened. Streaming needs care, since the first model may have streamed half a sentence before it died. RubyLLM can't take those chunks back, so the fallback starts a fresh assistant message, and `fallback.chunks_yielded?` tells you the user already saw something. Clear or mark the partial answer in your UI. The failed attempt is in the usage ledger, so `chat.cost` includes the half-answer you threw away.

`chat.cancel` stops a run at its next checkpoint by raising `RubyLLM::CancelledError`. The harder case is the stop button in a Rails app, where the stream runs in a job and the button lives in a different process:

```ruby
class ChatsController < ApplicationController
  def cancel
    current_user.chats.find(params[:id]).cancel # stops the job in another process
    head :no_content
  end
end

class ChatStreamJob < ApplicationJob
  def perform(chat_id)
    chat = SupportAgent.find(chat_id)
    chat.complete do |chunk|
      chat.messages.last&.broadcast_append_chunk(chunk.content) if chunk.content
    end
  rescue RubyLLM::CancelledError
    # broadcast a "stopped" state if your UI shows one
  end
end
```

`acts_as_chat` writes the request to the chat's `cancelled` column. The streaming job checks that column between chunks, at most once a second and outside the query cache, so it sees a write from another process without hammering your database. When it does, it clears the flag, raises, and removes the empty assistant row the stream created. The tokens already produced go into the ledger as a cancelled attempt. You don't need a Redis key or a pub/sub channel. Cancellation is cooperative: a tool in the middle of arbitrary Ruby code finishes before the next checkpoint notices, and closing a browser tab doesn't call `cancel` for you.

Provider errors come with subclasses like `RateLimitError` and `PaymentRequiredError`, all under `RubyLLM::Error`, and `error.response` keeps the HTTP response when there was one. Mistakes on your side (`ConfigurationError`, `ModelNotFoundError`, `PendingToolCallsError`) and `CancelledError` inherit straight from `StandardError`, so `rescue RubyLLM::Error` won't treat a user pressing stop as a provider failure.

Rescue blocks around every call site get copied and drift apart, so agents declare their policy once with `rescue_from`, the way Rails controllers do: `rescue_from RubyLLM::RateLimitError, RubyLLM::ServerError, with: :instrument_and_raise`. The semantics are `ActiveSupport::Rescuable`'s, and handlers cover `ask`, `complete`, `step`, and the other verbs. Handlers wrap agent instances: `SupportAgent.new.ask` goes through them, but `SupportAgent.find` and `SupportAgent.chat` hand you the record or chat itself, which doesn't. To get them on a persisted chat, use `SupportAgent.new(chat: Chat.find(id), persist_instructions: false)`.

Thanks to [@kieranklaassen](https://github.com/kieranklaassen) for asking for fallbacks ([#621](https://github.com/crmne/ruby_llm/issues/621)), [@sh1nj1](https://github.com/sh1nj1) for cancellable streams ([#607](https://github.com/crmne/ruby_llm/issues/607)), and [@skovy](https://github.com/skovy) for `rescue_from` ([#708](https://github.com/crmne/ruby_llm/issues/708)).

The [error handling guide](https://rubyllm.com/error-handling/) covers [model fallbacks](https://rubyllm.com/error-handling/#model-fallbacks) and the error hierarchy, [Rails streaming](https://rubyllm.com/rails-streaming/#cancelling-a-background-stream) covers the stop button, and the [agents guide](https://rubyllm.com/agents/#handling-errors-with-rescue_from) covers handlers.

## Prompt Caching

An agent sends the same system prompt, the same tool definitions, and the same 40-page contract on every turn. By turn twenty you've paid for that contract twenty times. Prompt caching lets the provider reuse a prefix it has already processed and bill the cache read at a fraction of normal input. RubyLLM 2.0 gives it one API:

```ruby
chat = RubyLLM.chat(model: "claude-sonnet-5").with_caching # one switch for every provider
chat.with_instructions(review_guidelines)

response = chat.ask "Review this migration.", with: "db/migrate/20261001_add_billing.rb"
response.tokens.cache_write # first request: the prefix goes into the cache
```

Anthropic wants `cache_control` markers on content blocks. Bedrock Converse wants cache points. OpenAI-compatible APIs take a `prompt_cache_key`. OpenAI and Gemini also cache on their own. `with_caching` translates to each, so none of that ends up in your code. Options set the TTL (`ttl: "1h"`), the cache key, and the mode. Agents get the same thing with `caching ttl: "1h"`, and `with_caching(false)` stops RubyLLM from sending cache controls, though it can't turn off caching a provider does on its own.

Automatic caching guesses where your reusable prefix ends. When you know, mark it:

```ruby
chat = RubyLLM.chat(model: "claude-sonnet-5").with_caching(ttl: "1h")

chat.with_instructions(analysis_prompt).cache_until_here # cache boundary
chat.add_message(role: :user, content: contract_text).cache_until_here

chat.ask "Does clause 14 conflict with the termination terms in clause 3?"
chat.ask "Which clauses mention liability caps?"
```

`cache_until_here` marks the latest message as a cache boundary, and RubyLLM sends your explicit boundaries plus automatic caching for the conversation growing after them. Anthropic, OpenRouter, Bedrock Converse, and selected OpenAI-compatible models understand boundaries; other providers keep their own caching behavior, so the same code runs everywhere. In Rails, the boundary is a `cache_until_here` column on the message, so it's replayed every time the chat is loaded, from a controller or a job days later.

Gemini and Vertex AI also let you create a cache as a resource with a name and an expiry, and point any number of chats at it. `RubyLLM.cache(content, model:, instructions:, ttl:)` creates one, `with_caching(id: cache)` uses it, and `RubyLLM::CachedContent.find(name, provider: :gemini)` finds it again from another process, with `renew` and `delete` to manage it. One Gemini rule: a request that uses a cache can't also send its own system instructions or tools, so put the instructions in the cache.

The providers still set the rules: minimum prefix length, how long the cache lives, which models support it. A hit is never guaranteed, and writing to a cache can cost more than plain input, so caching a prefix you never reuse costs more than not caching it. Check `tokens.cache_read` and `cost.cache_write` on your most repetitive workload. In Rails they're recorded per attempt, so whether caching paid off is a query.

Thanks to [@arunkumarry](https://github.com/arunkumarry), whose issue and pull request for Anthropic and Bedrock caching got this started.

The [prompt caching guide](https://rubyllm.com/prompt-caching/) has the per-provider option table and the full cache lifecycle.

## Batches

A user waiting for an answer needs an interactive request. An overnight job classifying ten thousand support tickets can wait. Anthropic, OpenAI, Gemini, Vertex AI, Bedrock, Azure, Mistral, xAI, OpenRouter, and Cohere all have batch APIs, and the major ones charge a lot less for it. Each has its own file format and steps for uploading, polling, and matching results. In RubyLLM 2.0, you write ordinary chats:

```ruby
chats = tickets.map do |ticket|
  RubyLLM.chat(model: "claude-haiku-4-5")
    .with_instructions("Classify this ticket as billing, bug, or feature request.")
    .ask_later(ticket.body)
end

batch = RubyLLM.batch(chats) # submits them all at once
batch.id # save this
```

`ask_later` stages the question without sending anything, and `RubyLLM.batch` submits all of them in one go, with instructions, history, tools, and schemas. Later, from any process, `RubyLLM::Batch.find(batch_id, provider: :anthropic)` and `batch.refresh.complete?` tell you whether it's done, and `batch.messages` comes back in submission order. A failed request is `nil` in `messages`, and `batch.statuses` tells you which slots succeeded, failed, or were cancelled, so you can resubmit the failures or finish them with `chat.complete`.

A batch generates one model turn. If the model asks for a tool, call `run_tools` on the chats and batch the ones that aren't complete yet. That's the [agentic loop](#the-agentic-loop) at batch prices.

When every input is a persisted chat, RubyLLM saves the batch in its own table, and a job can pick it up with nothing but the ID:

```ruby
class BatchPollJob < ApplicationJob
  def perform(batch_id)
    batch = RubyLLM::Batch.find(batch_id) # Rails already knows the provider
    return self.class.set(wait: 10.minutes).perform_later(batch_id) unless batch.refresh.complete?

    batch.messages
  end
end
```

RubyLLM restores the provider and the chats, then saves each answer through the same callbacks as a normal `ask`, so broadcasts, `after_create_commit` hooks, and usage records run as if the user had been waiting. Run the job twice and you still get one answer per chat. Embeddings batch too: `RubyLLM.embed_later` stages them for a catalog backfill, and `batch.results` returns the vectors in submission order.

`batch.cost` is a `RubyLLM::Cost`. RubyLLM uses the provider's reported cost when there is one, and otherwise prices each result at batch rates, which on Anthropic, OpenAI, Gemini, Bedrock, Azure, and Mistral means half the interactive price. The total stays `nil` until processing ends. On 2.0.0, batches from OpenAI reasoning models return a `nil` total because of how thinking tokens were priced; [Marc Köhlbrugge](https://github.com/marckohlbrugge) fixed that in 2.1.

Use one provider per batch. Anthropic and xAI accept mixed models; the others want one model per submission. Bedrock batches take no tools or structured output, Cohere's no structured output, and Bedrock and Vertex AI need a storage bucket for the batch files.

Thanks to [@marckohlbrugge](https://github.com/marckohlbrugge), [@thomaswitt](https://github.com/thomaswitt), [@toddkummer](https://github.com/toddkummer), and [@khasinski](https://github.com/khasinski), whose requests and feedback shaped this.

The [batch guide](https://rubyllm.com/batches/) has the full restriction table and setup.

## Owning the Transcript

A long conversation is full of things the user wants to keep and the model no longer needs to reread on every request. In 2.0 you can rewrite the history you send with `chat.messages = messages_for_model`. The setter takes `Message` objects, attribute hashes, or records that respond to `to_llm`, so you can summarize old turns, redact values, or drop a tangent, and the next request sends exactly what you put there. Your app knows what matters in a conversation better than a generic memory framework would.

This is practical because `message.content` is now a String or `nil`. Structured output is JSON text with a `parsed` reader, and files live on `message.attachments`. Conversations with tools need care: every tool call needs its result, and reasoning or provider-tool blocks have to stay with their message, so slicing the last four messages blindly can cut a call from its result, which providers reject.

On a Rails record, `messages=` is Active Record's association writer. For a temporary rewrite, use `chat_record.to_llm.messages = ...`. When the difference is permanent, for example the user sees everything and the model sees a redacted version, give RubyLLM its own association with `acts_as_chat messages: :llm_messages, message_class: "LlmMessage"`. Your UI renders `messages`, and RubyLLM persists and sends `llm_messages`.

Some providers can compact a long conversation themselves. `with_compaction(at: 100_000, instructions: "Keep every decision and every number.")` sets the input-token trigger and steers the summary. Anthropic and OpenAI Responses write an opaque compacted block that RubyLLM keeps and replays; OpenRouter drops messages from the middle once the context is full, with no threshold and no summary. `chat.compact` compacts on demand through OpenAI, Azure, and xAI Responses, and `chat.messages` keeps every original message. Either way, the summarization is billed work, and RubyLLM counts it.

When you need to see or change the final payload, there's `before_request`, and `render` shows the result without calling the model:

```ruby
chat = RubyLLM.chat(model: "gpt-5.6")
  .ask_later("Summarize the changes.")

chat.before_request do |payload|
  payload[:metadata] = { review_id: "review-42" }
end

chat.render[:metadata] # => { review_id: "review-42" }
chat.complete
```

The hook runs after all of RubyLLM's formatting and provider-option merging. It speaks the selected protocol's wire format, so a hook written for OpenAI needs revisiting if you move to Anthropic. `render` also lets you test request shaping without network access or an API key. For a fixed field like this one, `with_provider_options(metadata: ...)` is simpler.

Finish reasons are now normalized symbols: `:stop`, `:max_tokens`, `:tool_calls`, or `:content_filter`, with `stopped?`, `max_tokens?`, `tool_call_stop?`, and `content_filtered?`. A truncated answer is `:max_tokens` whether the provider said `length`, `max_tokens`, or `MAX_TOKENS`, and Anthropic's `end_turn`, Gemini's `STOP`, and the Responses API's `completed` are all `:stop`. Reasons RubyLLM doesn't map, like Anthropic's `pause_turn`, come through as symbols in the provider's spelling. On Anthropic, `:max_tokens` also covers a full context window, so check `with_max_output_tokens` first and the transcript second.

Thanks to [@mnort9](https://github.com/mnort9) and [@marksweston](https://github.com/marksweston) for pushing on transcript control, [@fvaleye](https://github.com/fvaleye) for compaction, and [@trevorturk](https://github.com/trevorturk) and [@losingle](https://github.com/losingle) for finish reasons.

Details are in the [request control guide](https://rubyllm.com/chat-request-control/) and the [persistence guide](https://rubyllm.com/rails-persistence/#separate-user-and-llm-transcripts).

## Providers and Protocols

Cohere, Ollama Cloud, ElevenLabs, and Deepgram join the list, bringing the total to seventeen.

The bigger change is how much of each provider RubyLLM covers. I audited forty shared features across all seventeen providers, comparing 1.16 with 2.0. Built-in support went from 170 provider-feature pairs to 395, out of the 405 those providers offer. Those count provider-feature pairs, not distinct features or every endpoint: how much of each provider's API you can reach from the same Ruby code. The [coverage matrix](https://rubyllm.com/provider-coverage/) shows every cell, with sources and the gaps that remain.

Under the hood, providers and protocols are now separate. A provider is the service you connect to; a protocol is the API it speaks. OpenAI defaults to the Responses API, and Vertex AI and Bedrock route each model to the API it actually speaks:

```ruby
RubyLLM.chat(model: 'gpt-5.4')                              # OpenAI Responses API
RubyLLM.chat(model: 'gpt-5.4', protocol: :chat_completions) # same model, old API
RubyLLM.chat(model: 'claude-opus-4-6', provider: :vertexai) # Vertex AI, Anthropic protocol
```

Responses gives RubyLLM access to encrypted reasoning, provider-run tools, and newer OpenAI features. If you depend on Chat Completions, pass `protocol:` per chat or set `config.openai_protocol = :chat_completions`. Thinking has one API too: `with_thinking` uses the model's defaults, or pass `effort:` or `budget:`, and read it back through `response.thinking` ([guide](https://rubyllm.com/thinking/)).

A new provider can be a single small file. If yours is missing, generate a gem for it:

```bash
ruby_llm provider-gem Acme --api-base https://api.acme.ai/v1
```

I covered the design in [Providers, Protocols, and Provider Gems](/rubyllm-2-0-providers-and-protocols/).

## Rails Table Ownership

A fresh 1.x install put four models in your app: `Chat`, `Message`, `Model`, and `ToolCall`. In 2.0 it's two:

```ruby
# app/models/chat.rb
class Chat < ApplicationRecord
  acts_as_chat
  belongs_to :user
end

# app/models/message.rb
class Message < ApplicationRecord
  acts_as_message
  has_many_attached :attachments
end
```

Chats and messages are your product's conversations, so they stay yours: users, scopes, authorization, titles, retention. Everything else lives in tables RubyLLM owns under a `ruby_llm_` prefix: `ruby_llm_models` for the registry, `ruby_llm_tool_calls` for tool requests, approval decisions, and links to their results, `ruby_llm_usages` with one row per provider attempt, and `ruby_llm_batches`. You read them through the same API as in plain Ruby: `RubyLLM.models`, `message.tool_calls`, `message.tokens`, `chat.cost`, and `RubyLLM::Batch.find`.

In 1.x the generator copied `Model` and `ToolCall` into your app, and from then on they were your problem. When RubyLLM needed to store something new about a tool call, you had to update a class you didn't write and never called directly. 2.0 stores a lot more: approval decisions, a usage row per attempt, batch state. Shipping that as "please update these four files" would have made a painful upgrade, and the next feature would have needed another one.

Rails has a precedent. You don't have an `ActiveStorageBlob` in `app/models`; Active Storage owns its records and you use them through `has_many_attached`. RubyLLM's tables are ordinary tables created by ordinary migrations, not an engine. The record classes behind them are internal and the readers are public, so RubyLLM can change its storage without your app noticing. Every feature in this post that needed new columns got them without asking you to maintain another model.

Usage rows also make your own reporting plain Active Record: `has_many :ruby_llm_usages, through: :chats` on `User`, then `current_user.ruby_llm_usages.sum(:total_cost)`. Keep the `nil` caveat from [cost and usage](#cost-and-usage) in mind.

If you put your own columns on the old `models` table, like availability flags, default models, or admin pricing, those are product features. Give them a table of your own keyed by provider and model ID. The [upgrade](#upgrading) keeps the columns, but RubyLLM won't maintain them.

The [Rails persistence guide](https://rubyllm.com/rails-persistence/) shows every association and reader.

## Prompt Templates

A two-page prompt inside a Ruby heredoc is hard to work with. The prompt changes often, the class around it rarely does, and every diff is paragraphs of English between `def` and `end`. Prompts are templates, so in 2.0 they get a directory the way views do, `app/prompts`, and `RubyLLM.render_prompt` renders them:

```erb
<%# app/prompts/support/instructions.txt.erb %>
You are a support assistant for <%= product_name %>.
The current customer is <%= customer_name %>.
Answer with concise, practical steps.
```

```ruby
instructions = RubyLLM.render_prompt(
  "support/instructions",
  product_name: "BillingHub",
  customer_name: current_user.name
)
```

It reads an ERB file, renders it with your locals, and returns a String. It doesn't call a model or create a chat, so use the string as instructions, as a user message, or as input to `RubyLLM.embed`. Lookup is relative to `Rails.root` in Rails and the current directory in plain Ruby. A missing local makes ERB raise, and a missing file raises `RubyLLM::PromptNotFoundError`, so you can test every prompt without an API key. Prompt files are code, since ERB runs Ruby: keep prompt names in your code and pass user input only as locals.

Agents have used `app/prompts` since [1.12](/rubyllm-1-12-agents/), and now share the same renderer. In 2.0 a named agent picks up its prompt by convention: `WorkAssistant` reads `app/prompts/work_assistant/instructions.txt.erb` if it exists, and starts without instructions if it doesn't. When the prompt is required, declare `instructions { prompt("instructions") }` and a missing file raises. Locals can be lambdas that run with access to the chat and the agent's inputs, and instructions can stack: `instructions append: true, persist: false` adds text, like today's date, that stays out of the saved transcript.

A Rails-backed agent saves its instructions when `create!` makes the chat, and `WorkAssistant.find(id)` applies the current configuration without rewriting what's saved. When old conversations should pick up new wording, `WorkAssistant.sync_instructions(chat)` re-renders and saves them (it was `sync_instructions!` in 1.x). A bare `instructions` call used to require the conventional prompt; in 2.0 it's only the reader.

Rails engines can ship prompts. `RubyLLM::Prompt.roots` is an ordered list of prompt directories, like Action View's view paths, with the app's `app/prompts` always first. An engine appends its own directory, and the host app overrides any engine prompt by creating the same path in `app/prompts`. This came from [@adrianthedev](https://github.com/adrianthedev). Thanks to [@kryzhovnik](https://github.com/kryzhovnik), who extracted the renderer out of Agent's private methods, which made this possible.

The [prompt rendering guide](https://rubyllm.com/prompt-rendering/) has the details, and the [agents guide](https://rubyllm.com/agents/) covers the class-based conventions.

## Workflows and Instrumentation

A research agent calls a model, searches the web a few times, and hands its notes to a writing agent. The Ruby is short, but the logs show dozens of separate events with nothing tying them to the article they produced. 2.0 lets you name the work:

```ruby
RubyLLM.workflow("Write article", id: "article-42") do |workflow| # tags every event inside
  notes = workflow.step("Research") do
    ResearchAgent.new.ask(topic).content
  end

  workflow.step("Draft") do
    WriterAgent.new.ask(notes).content
  end
end
```

Every RubyLLM event inside the block carries `workflow_id` and `workflow_name`, and inside a step, `workflow_step_id` and `workflow_step_name`. Model calls, tool calls, and usage rows for retries are all tagged. The blocks return their normal values, so `notes` is a String.

Many AI frameworks come with a graph DSL: nodes, edges, a state object, and a runtime that executes it. In Ruby, a sequence is method calls, a branch is a `case`, and a retry is `retry`. What's missing is a way to see that control flow afterwards, and that's all `RubyLLM.workflow` does. It doesn't persist progress, schedule anything, or retry. For durability, use the [loop's verbs](#the-agentic-loop) and your job queue.

1.16 introduced instrumentation with five events. 2.0 has twenty: every model operation, per-attempt usage, and the workflow and step wrappers. In Rails they go through `ActiveSupport::Notifications`, so you subscribe the way you'd subscribe to `sql.active_record`. Subscribe to `usage.ruby_llm` and group by `workflow_id` to get the cost of a piece of work, failed attempts included. Outside Rails, set `config.instrumenter` to anything that responds to `instrument(name, payload)` and yields. If you wrote subscribers for 1.16, tokens moved under `payload[:tokens]`.

One-shot operations take `metadata:`, which lands in `payload[:metadata]` and is never sent to the provider. Workflows take it too, and nested events get it as `workflow_metadata`. Steps and workflows nest, recording parent IDs, so a subscriber can rebuild the execution tree of a run. Context lasts exactly as long as the block, so it doesn't follow your work into a job that runs tomorrow; save the ID and open the workflow again with the same `id:`. When you run work concurrently yourself, open the step inside each task.

Payloads include message content, tool arguments, and provider responses, so export those only where your policy allows it. The [instrumentation guide](https://rubyllm.com/instrumentation/) lists every event and payload field.

## Coding Assistant Skill

The gem ships a RubyLLM skill for coding assistants, maintained alongside the code and guides. Install the copy that matches your bundle:

```bash
npx skills add "$(bundle show ruby_llm)" --skill rubyllm
```

Your assistant stops writing 1.x method names from its training data and stops rebuilding things RubyLLM already does. [Setup details](https://rubyllm.com/ai-coding-assistants/).

## Upgrading

There are two jobs: update your Ruby code, and, if you use Rails persistence, migrate your stored records. Plain Ruby apps only do the first. Your chats and messages keep their IDs and relationships, nothing is deleted until you run cleanup, and every migration phase can be retried.

**Start on 1.16.** If you're on something older, step through the minor releases first, because the generator expects the schema 1.16 produced. While still on 1.16, set `config.deprecation_behavior = :raise` in your test environment and fix whatever breaks, and if you still have `config.use_new_acts_as = false`, switch to the association-based `acts_as` API, the only one 2.0 has.

**Pin the 2.0 series.** Then run `bundle update ruby_llm`:

```ruby
gem "ruby_llm", "~> 2.0.0"
```

The 1.16 upgrade generator, its migration helpers, and the copy-mode tasks ship with 2.0 only. From 2.1 on, each release carries just the upgrade from the release before it. Finish this one, cleanup included, before you move to 2.1.

**Update your calls.** Most changes are renames, because each concept now has one name: `with_tool` becomes `with_tools`; tool `desc`, `param`, and `params` become `description`, `parameter`, and `parameters`; `with_params` becomes `with_provider_options`; `response.input_tokens` becomes `response.tokens.input`; `on_new_message` and `on_tool_call` become `before_message` and `before_tool_call`; `create_user_message` becomes `ask_later`; `RubyLLM::Schema` becomes `Schematist::Schema`. Rails controller `params` have nothing to do with the `params:` rename, so change RubyLLM call sites one by one rather than with search-and-replace. If you tried a release candidate, rename `with_server_tools` to `with_provider_tools`.

A few behavior changes matter more than names. Finish reasons are Symbols. Switches like `with_thinking`, `with_caching`, `with_citations`, and `with_compaction` take `false` to turn off, and value setters like `with_temperature` take `nil` to reset. The new `before_` and `after_` callbacks stack, where the old `on_*` ones replaced each other. OpenAI uses the Responses API, so Chat Completions-only options like `response_format` need the Responses shape or `config.openai_protocol = :chat_completions`. The temperature you set is the temperature sent; 1.x sometimes rewrote it for certain models.

**Choose rename or copy.** Rename, the default, moves your existing model and tool-call tables into RubyLLM's ownership. It's fast, needs AI activity paused for all three phases, and going back means restoring your backup with the matching app build. Copy (`--mode copy`) builds the 2.0 tables next to the 1.16 ones, runs preparation and backfill while 1.16 keeps serving traffic, and keeps a supported route back to 1.16 with `ruby_llm:upgrade:rollback` and `ruby_llm:upgrade:resume`. On 100,000 synthetic chats with a million messages on PostgreSQL:

| Mode | Total migration time | AI downtime |
|---|---:|---:|
| Rename | 20 s | 20 s |
| Copy | 136 s | 4 s |

Those are medians from my [migration benchmark](https://github.com/crmne/ruby_llm_migration_bench), not a promise about your database. Copy mode generates a compatibility concern and initializer that must run in both builds, and its guards don't cover `update_columns`, bulk SQL, direct deletes, or attachment purges. Its migrations also need a direct database connection or a session-mode pool. If you don't need the way back or the shorter pause, use rename.

**Generate and rehearse.** `bin/rails generate ruby_llm:upgrade` (add `--mode copy` and any custom class mappings) writes three migrations: prepare, backfill, and finish. Backfill works in batches of 10,000 messages and saves its progress, so a rerun skips finished batches and never counts historical usage twice. Rehearse on a recent copy of production with the exact build you'll deploy, time it, and check message counts, tool-result links, and token counts. Take a backup you've actually restored before.

**Run it.** For rename, finish or cancel conversations waiting on tool results, stop everything that touches AI, deploy 2.0, and run `bin/rails db:migrate` and `bin/rails ruby_llm:load_models`. Copy mode deploys its compatibility code to 1.16 first, runs prepare and backfill from the 2.0 build while 1.16 serves, then pauses only for finish. If your deploy's migration step has a timeout, generate the phases one at a time with `--phase`.

**Delete two models.** Remove `Model` and `ToolCall` with their `acts_as_model` and `acts_as_tool_call` declarations, drop `model:` and `tool_calls:` from the remaining macros, and remove `config.model_registry_class` and `config.use_new_acts_as`. Don't drop the old tables yourself. The backfill creates one usage row per historical response with recorded usage, but it can't recover retries or prices 1.16 never recorded.

**Clean up later.** Once you trust 2.0 in production, generate `--phase cleanup` in a later release. In copy mode, run `bin/rails ruby_llm:upgrade:finalize` first, which closes the way back. Then delete the 2.0 upgrade migrations from `db/migrate`.

The migrations have no `down`, and changing the gem version back does not roll back the database. The [complete 2.0 upgrade guide](https://github.com/crmne/ruby_llm/blob/v2.0.0/docs/_reference/upgrading.md) has every rename, the copy-mode procedure step by step, incomplete tool calls, and custom class names.

## Get It

```bash
bundle add ruby_llm
```

[What's New in 2.0](https://rubyllm.com/whats-new-in-2-0/) has working examples of everything above, the [upgrade guide](https://github.com/crmne/ruby_llm/blob/v2.0.0/docs/_reference/upgrading.md) walks a Rails app from 1.16, and the [documentation](https://rubyllm.com) covers the rest.

Thanks to everyone who tested the release candidates, filed issues, and sent code, including 22 people who made their first contribution during this cycle.

RubyLLM 2.1 comes out on October 8, at Deccan Queen on Rails in Pune.
