RubyLLM 2.0: Agents That Wait, and Everything Else Your Providers Can Do

RubyLLM 2.0: Agents That Wait, and Everything Else Your Providers Can Do

RubyLLM 2.0 is the biggest release since 1.0. It starts with tool approval:

class RefundOrder < RubyLLM::Tool
  description "Refunds an order"
  parameter :order_id, description: "ID of the order to refund"
  requires_approval # that's all it takes

  def execute(order_id:, tool_call: nil)
    Refunds.issue(order_id:, idempotency_key: tool_call.id)
  end
end

class SupportAgent < RubyLLM::Agent
  chat_model Chat
  tools RefundOrder
end

chat = SupportAgent.create!
chat.ask "I was charged twice for order 42." # stops before refunding

# later, when the user approves
chat = SupportAgent.find(chat.id)
chat.approve(chat.pending_approvals.first)
chat.complete # issues the refund and answers

The pending call is saved in your database, so the approval can come tomorrow, from another process, after two deploys.

That’s just one piece. 2.0 also generates video and speech, reads scanned documents, reranks search results, runs provider-hosted tools with typed citations, accounts for every provider attempt (including the ones that failed), and works across seventeen providers.

This post covers all of it, one section per feature. Each section links to the guide with the full details.

Tool Approval and Durable Agents

Refunds, deletions, and customer emails shouldn’t happen just because a model asked. A person should see the exact call and its arguments first. Declare requires_approval on the tool, and the loop parks instead of running it:

chat = RubyLLM.chat.with_tools(RefundOrder)
chat.ask "Refund order 42."

chat.awaiting_approval? # => true, and nothing has executed
call = chat.pending_approvals.first
call.name               # => "refund_order"
call.arguments          # => {"order_id" => "42"}

chat.approve(call)
chat.complete           # runs the refund, then the model answers

ask runs the loop until there’s nothing left it’s allowed to do, then returns. No thread blocks waiting for a human: the chat records the request, and your code decides.

chat.deny(call) skips the call. The model gets a tool result that reads The user denied the refund_order tool call. and can explain, offer something else, or ask what you’d prefer. Nothing raises. Tools that don’t need approval still run in the same round, so if the model asks to look up the order and refund it, the lookup happens and the refund waits.

A decision has three values: true (approved), false (denied), and nil (still waiting). If your app already stores approvals, give the tool a resolver block, requires_approval { |call| ... }, that returns one of those three. The resolver is then the source of truth. RubyLLM may call it several times while a call waits, including after a restart, so keep it a read.

While a call is unanswered, ask and ask_later raise RubyLLM::PendingToolCallsError. Providers reject a transcript with a tool call that has no result, so RubyLLM raises before sending it.

Pausing in memory is easy. The harder case is a human who takes twenty minutes to answer while you deploy twice. With acts_as_chat, the tool call and its decision are database rows, so the approval can come from anywhere:

class ApprovalsController < ApplicationController
  def create
    chat = Chat.find(params[:chat_id])
    chat.approve(params[:tool_call_id]) # or chat.deny(...)
    CompleteJob.perform_later(chat.id)
  end
end

class CompleteJob < ApplicationJob
  def perform(chat_id)
    SupportAgent.find(chat_id).complete # resumes where it stopped
  end
end

SupportAgent.find reloads the transcript and reapplies the agent’s tools and configuration. Load through the agent, not Chat.find, when you resume: a bare record has the transcript but no tools, so it doesn’t know RefundOrder needs approval and has nothing to run once it’s approved. Tool calls that already have saved results are skipped. Render pending_approvals as cards showing each call’s name and arguments. There’s no workflow engine and no state machine to serialize: the transcript holds the state.

One rule comes with it: tools that write should be idempotent. A process can die after a tool’s side effect succeeds and before its result is saved, and the next job runs the approved call again. That’s why the refund passes tool_call.id to the payment provider as an idempotency key. For writes in your own database, a unique constraint on the call ID does the same job.

Approval also covers provider-run tools. When OpenAI or Azure calls a tool on a remote MCP server, the provider can pause before the call, and it lands in the same pending_approvals list with call.remote? set. See provider tools.

Thanks to @jondavidschober for the issue that started this (#503).

The tool approval guide and durable agents guide cover the full flow, including ActiveJob Continuations for surviving deploys mid-run.

The Agentic Loop

ask still runs the whole agentic loop for you. When you want control, 2.0 breaks it into verbs:

chat = RubyLLM.chat.with_tools(SearchDocs).ask_later("How do I configure webhooks?")
chat.step until chat.complete? || chat.awaiting_approval?

ask_later stages a message. generate makes one model call. run_tools runs pending tools. step picks the next move. complete keeps going until there’s an answer or a decision to make, and ask is ask_later followed by complete. cancel stops a run from another thread, or, in Rails, from another process through the database.

Each verb decides what to do next by reading the persisted messages, so the loop doesn’t need to live in one process. You can run one step per job, each with its own retry boundary. run_tools skips calls whose results have been saved, so a process that dies after one result out of three resumes with the remaining two. With these verbs, iteration budgets and handoffs between agents each take a few lines of Ruby.

halt and RubyLLM::Tool::Halt are gone. In 1.x a tool could end the conversation from its return value. In 2.0 your loop decides when to stop, and chat.with_tool_options(calls: :one) limits the model to one tool call per response.

I wrote about this in detail in The Agentic Loop, Exposed. The agentic workflows guide has the patterns built on these verbs.

New Operations

1.x could chat, embed, paint, transcribe, and moderate. 2.0 adds video, speech, OCR, reranking, multimodal embeddings, and tokenization:

RubyLLM.animate("A red panda typing Ruby code").save("panda.mp4")
RubyLLM.speak("Welcome to the Ruby study group.").save("welcome.mp3")
RubyLLM.ocr("scanned-contract.pdf").markdown
RubyLLM.rerank("How do I reset my password?", documents, model: "rerank-v3.5")
RubyLLM.embed("A red panda", with: "panda.jpg", model: "gemini-embedding-2")
RubyLLM.tokenize("Hello, Ruby!", model: "grok-4.3").ids

None of them needs a chat. Results are typed, and save and to_blob work on images, speech, video, and downloaded files.

Video. animate submits a video job, polls until the provider finishes, and returns a RubyLLM::Video. Pass a still image with with:, as in chats, to animate it, and extend: to continue a video you generated. animate_later returns a RubyLLM::VideoJob right away. Call refresh on it from any process; done? is true once the job finishes either way, completed? only when you got a video, and wait runs the polling loop for you. Gemini, Vertex AI, xAI, Azure, Bedrock, OpenRouter, ElevenLabs, and GPUStack all sit behind the same call.

OCR. RubyLLM.ocr returns markdown per page, plus the images and tables the provider found and its raw page data, without involving a chat model. Mistral Document AI reads PDFs, office documents, and images; Cohere Parse reads one image per request. pages: takes zero-based page indexes, because every OCR API has that concept. Table formats differ between providers, so options like table_format go in provider_options:, in the provider’s own words.

Reranking. Embeddings are good at finding fifty documents that might answer a question. They’re worse at deciding which five actually do. A reranker reads the query against each candidate and sorts them, and each result’s index maps it back to your records:

query_vector = RubyLLM.embed(query).vectors
candidates = Article.nearest_neighbors(:embedding, query_vector, distance: :cosine).limit(50).to_a
rerank = RubyLLM.rerank(query, candidates.map(&:body), model: "rerank-v3.5", top_n: 5) # keeps the best 5

best = rerank.results.map { |result| candidates[result.index] }

model: is required, because reranker catalogs differ per provider and there’s no sensible default. Scores belong to the model that produced them: fine for ordering and for a cutoff you tune on your own data, meaningless to compare across rerankers. Cohere, Bedrock, Vertex AI, Azure, OpenRouter, and GPUStack all answer rerank.

Tokens. RubyLLM.tokenize returns the token IDs for exactly the string you pass, on xAI, Cohere, and GPUStack. Usually you want the size of a whole request, so ask the chat: chat.count_tokens("Summarize this contract.") counts history, instructions, function tools, schema, thinking settings, and supported attachments through the provider’s own counting endpoint, without adding the question to the conversation. RubyLLM.count_tokens(text, model:) does the same for a lone prompt. Anthropic, OpenAI, Gemini, Vertex AI, and Bedrock count requests, and anything unsupported raises RubyLLM::Error instead of guessing.

Counting endpoints don’t receive provider tools, provider_options, compaction, or before_request edits, so a chat using those can send something slightly different from what you counted. Neither count predicts the output; response.tokens has the actual counts afterwards.

Each one has a guide: video, OCR, reranking, embeddings, and tokenization.

Speech and Transcription

RubyLLM has had transcription since 1.x. 2.0 adds text to speech with RubyLLM.speak, which returns a RubyLLM::Speech with model, voice, format, mime_type, to_blob, and save. With both, a Rails action can answer a voice message with audio:

def create
  audio = params[:audio]
  return head :bad_request unless audio.is_a?(ActionDispatch::Http::UploadedFile)

  question = RubyLLM.transcribe(audio).text
  answer = Current.user.support_chat.ask(question).content

  speech = RubyLLM.speak(answer) # text to speech
  send_data speech.to_blob, type: speech.mime_type, disposition: "inline"
end

Check that the upload is actually a file: a bare string parameter would make RubyLLM read a path or fetch a URL you didn’t intend. The middle step is a normal chat, so it can use tools, approvals, and history. The three steps are ordinary sequential requests, not a realtime audio stream.

OpenAI, Gemini, ElevenLabs, Deepgram, Mistral, xAI, Azure, Vertex AI, OpenRouter, and GPUStack all answer to speak. model:, voice:, and format: are keywords because every provider has them. Voices are provider-specific: OpenAI and Gemini use names, ElevenLabs uses voice IDs from your voice library, and Deepgram bakes the voice into the model name. Everything else, like OpenAI’s instructions and speed, ElevenLabs’ voice_settings, goes in provider_options: in the provider’s own words. I pass each provider’s options through rather than invent a lowest-common-denominator emotion: parameter that would cover only part of what each one supports.

Pass a block and the audio streams as SpeechChunks, while the call still returns the complete Speech, so you can play first and save after. Two Gemini caveats: its speech endpoint doesn’t stream, and asking it to raises instead of quietly buffering; and it returns raw PCM, so speech.format is "pcm" and naming the file .wav doesn’t make it WAV. Run it through ffmpeg when you need a container.

Transcription gained streaming and shared controls for speakers and timestamps:

transcription = RubyLLM.transcribe(
  "standup.wav",
  model: "gemini-3.5-transcribe",
  speaker_names: [], # speaker labels
  timestamps: :word # per-word timing
)

transcription.words.each do |word|
  puts "#{word['speaker']} at #{word['start']}s: #{word['word']}"
end

Speaker labels now work beyond OpenAI, and OpenAI’s diarization model can still match voices to names from short reference clips with speaker_references:. Pass a block to stream: delta? marks committed text, partial? a tentative guess the next chunk replaces, and segment? speaker and timing data. Deepgram, ElevenLabs, xAI, and Google’s live models stream over WebSockets with the optional websocket-driver gem; OpenAI, Mistral, and others stream over plain HTTP. Either way, the input is a recording you already have. format: asks for the provider’s transcript format, such as OpenAI’s "srt".

The text to speech guide and transcription guide have every model, voice, and format.

Provider Tools

Providers now run tools on their own infrastructure: web search, code execution, file search, remote MCP servers. With with_provider_tools, web search needs no search client or extra key:

chat = RubyLLM.chat(model: "claude-sonnet-5").with_provider_tools(:web_search) # runs on Anthropic's side
response = chat.ask "What changed in the latest Ruby release?"

response.citations.each { |citation| puts "#{citation.title}: #{citation.url}" }

Your RubyLLM tools run in your process, provider tools on the provider’s side, and one chat can use both: chat.with_tools(Weather).with_provider_tools(:web_search, :code_execution). The aliases make provider tools portable. :web_search becomes Anthropic’s versioned search tool, OpenAI’s Responses search tool, Gemini’s Google Search grounding, and the equivalent on other providers, so you can switch models without changing the code. The full set is :web_search, :web_fetch (or :url_context), :x_search, :code_execution, :file_search, :image_generation, :apply_patch, and :mcp.

Options use the provider’s own vocabulary, like web_search: { allowed_domains: ["ruby-lang.org"], max_uses: 3 }, because a lowest-common-denominator filter language would lose what each provider supports. The alias works across providers; the options are specific to one. Providers also ship new tools faster than any library can wrap them, so a positional Hash goes into the payload as is. RubyLLM handles the response blocks the protocol returns, so you need a definition the endpoint accepts, not a new gem release.

With :mcp, the provider connects to a remote MCP server, lists its tools, and calls them, and your process never opens a connection to it. On OpenAI and Azure Responses, the provider can stop and ask before each call, and those requests land in the same approval flow as your Ruby tools:

chat = RubyLLM.chat(model: "gpt-5.6")
  .with_provider_tools(mcp: {
    name: "docs",
    url: "https://learn.microsoft.com/api/mcp",
    allowed_tools: ["microsoft_docs_search"],
    require_approval: "always" # the provider asks first
  })

chat.ask "Search Microsoft documentation for Azure Blob Storage."
call = chat.pending_approvals.first
call.remote? # => true

chat.approve(call) # or chat.deny(call)
chat.complete

Tool activity shows up on the response: response.server_tool_calls gives each call’s name, input, result, and the provider’s raw block. Search results are the same Citation objects you get from documents, and images from :image_generation or files from :code_execution come back as response.attachments. Providers often bill tool uses on top of tokens, so response.tokens.server_tool_use gives you the per-use counters to price them.

Some providers need their tool blocks replayed on later turns, and Anthropic rejects a conversation that drops them. RubyLLM keeps them on the message and sends them back. In Rails, the server_tool_calls and raw_content columns keep them across requests, and an agent can declare provider_tools :web_search, :code_execution so they come back with the chat.

Availability depends on the provider, model, and protocol. Ask for an alias a provider doesn’t have and you get RubyLLM::UnsupportedServerToolError before any request, listing the ones it does. A few tools need a non-default protocol: Mistral’s hosted search and code execution want protocol: :conversations, Gemini’s remote MCP wants protocol: :interactions, and OpenRouter’s hosted shell and MCP want protocol: :responses. Not every provider can pause for MCP approval either: Anthropic, Gemini Interactions, and xAI run allowed MCP tools immediately, so choose which tools to allow.

The provider tools guide has the per-provider tables, file search setup, and MCP connection details.

Citations

When a model tells your user “The contract can be terminated with 30 days’ notice,” the user needs to know where that comes from and which page to check:

chat = RubyLLM.chat(model: "claude-sonnet-5").with_citations # document citations
response = chat.ask "What are the termination conditions?", with: "contract.pdf"

response.citations.each do |citation|
  puts "p. #{citation.start_page}: #{citation.cited_text}"
end

with_citations turns on document citations for Anthropic, Cohere, and Claude on Bedrock, and with_citations(false) turns them off. Web search citations need no setting: turn on the provider’s search tool and they show up, and Perplexity’s Sonar models search on every request.

Every provider has its own citation format. Anthropic returns citation blocks. Gemini returns grounding metadata with byte offsets. OpenAI hangs annotations off the response text. Perplexity sends a list of URLs on every chunk. If you build footnotes against each one directly, you write the footnote feature once per provider. RubyLLM turns all of them into RubyLLM::Citation, with url, title, cited_text (the quoted passage), text (the part of the answer it supports), start_index and end_index, source_id, source_index, and start_page and end_page. Whatever the provider didn’t report is nil; a web citation often has a URL and nothing else.

Positions are character offsets into response.content, ready to slice. Gemini reports byte offsets, which would drift the moment your answer contains an accented letter, so RubyLLM converts them and your footnote lands after “Café” and not in the middle of it.

Citations matter most in RAG, which is also where they usually get lost: your tool returns a JSON blob, and the model paraphrases it with no trail back. Return RubyLLM::SearchResults instead:

class KnowledgeBase < RubyLLM::Tool
  description "Searches the company knowledge base"
  parameter :query, description: "What to look for"

  def execute(query:)
    docs = Document.search(query)
    return "No matching documentation found." if docs.empty?

    RubyLLM::SearchResults.new( # citable results
      *docs.map { |doc| { title: doc.title, url: doc.url, text: doc.body } }
    )
  end
end

Anthropic, Cohere, and Bedrock receive these as their native citable search results, including after a Rails chat is reloaded from the database, and response.citations.first.url is the doc.url you returned. Other providers get the results as JSON text, which they can read but not formally cite.

Citations arrive on streaming chunks, and the final message collects them all, deduplicated. Some providers only send citations at the end, so build the finished footnotes from the final message. With acts_as_message, they’re saved as JSON and come back as Citation objects. One limit: Anthropic won’t combine document citations with with_schema, so a single request gets either citations or structured output.

The citations guide has a footnote renderer you can copy.

Files and Attachments

If a user asks five questions about the same 40 MB PDF, there’s no reason to send those 40 MB five times. Upload it once:

manual = RubyLLM.upload("manual.pdf", provider: :anthropic) # upload once

chat = RubyLLM.chat(model: "claude-sonnet-5")
chat.ask "Summarize the manual.", with: manual
chat.ask "What does the warranty cover?"

RubyLLM.chat(model: "claude-sonnet-5").ask "Which page explains the reset button?", with: manual

The requests point at the provider’s copy, and with: takes an uploaded file like a path or a URL. This saves the transfer, but you still pay for the document’s tokens on every request. Prompt caching is what makes re-reading it cheap.

RubyLLM.upload takes a path, an IO, or a RubyLLM::Attachment, and returns a RubyLLM::UploadedFile with its id, provider, filename, byte_size, mime_type, and expires_at. A file ID only works with the provider that issued it, so if you store IDs, store the provider next to them and use RubyLLM::UploadedFile.find(id, provider:). RubyLLM.download(file_id, provider: :openai) brings files back with the same save and to_blob as other media. Twelve providers have file APIs behind this, including S3 for Bedrock and Cloud Storage for Vertex AI. Fewer can reference a stored file from a chat, and download rules are the provider’s: Anthropic and OpenRouter, for example, only let you download files their tools created.

Most of the time you won’t call upload at all. Attach the file as usual, and if it’s too big to send inline, RubyLLM uploads it and sends a reference instead. Gemini switches over at 20 MB (7 MB on Vertex AI), Anthropic at 24 MB, OpenAI at 50 MB for PDFs and documents. The upload is remembered on the attachment per provider and per set of credentials, so follow-ups reuse it, a provider switch uploads once there, and an expired file is uploaded again. Your local history keeps the original file. The memo lives in memory, though, so a Rails chat reloaded in another process uploads its stored file again. config.auto_upload_large_files = false turns this off.

Tools can return files too, so a chart tool can return the chart:

class RevenueChart < RubyLLM::Tool
  description "Renders a revenue chart for a quarter"
  parameter :quarter, description: "Quarter, like 2026-Q3"

  def execute(quarter:)
    path = Charts.revenue(quarter).render_png
    ["Revenue chart for #{quarter}", RubyLLM::Attachment.new(path)] # text plus a file
  end
end

Strings become the tool result’s text and attachments become its files. A browser tool can return a screenshot for a vision model to look at. Anthropic and Bedrock take tool files inside the tool result, Gemini 3 inside the function response, and OpenAI, whose tool results are text-only, gets them in a user message right after the result. You write the return value once. If a provider can’t take that file type at all, you get RubyLLM::UnsupportedAttachmentError rather than a model answering about a file it never saw.

The files guide covers expiration, provider limits, and storage setup, and the attachments guide covers everything with: accepts.

Cost and Usage

Chat history doesn’t match what you pay for. A retry that failed, a fallback model that took over, and a stream the user cancelled halfway all cost tokens, and none of them leave an assistant message behind.

2.0 records usage per provider attempt, because providers bill attempts and one message can take several:

chat = RubyLLM.chat.with_fallbacks("claude-sonnet-5")
response = chat.ask "Explain Ruby fibers in one paragraph."

response.tokens.input
response.cost.total
chat.cost.total # every attempt, including the ones that failed

response.tokens and response.cost add up every attempt that produced that response: retries, fallbacks, and the one that succeeded. chat.tokens and chat.cost add up every attempt the chat ever made, including cancelled ones with no message to show for it. One-shot operations like embed and transcribe return the same objects.

Providers disagree on whether cached tokens are part of the input count, so RubyLLM normalizes the buckets before pricing. tokens.input is ordinary input excluding cache reads and writes, and output, cache_read, cache_write, and thinking mean the same thing on every provider. When a provider bills thinking as output, it’s already in tokens.output, so don’t add it twice. cost mirrors the buckets and adds total.

When usage or pricing is unknown, the total is nil rather than zero, so a request you can’t price doesn’t look free. That covers a model missing from the pricing registry and an attempt that may have been billed but never reported usage: a timeout, a server error, a stream that died early. An attempt that provably wasn’t billed, like a refused connection or a 4xx rejection before the model ran, records zero. Provider-reported charges from OpenRouter or xAI are kept in tokens.reported_cost and win over estimates. Hosted tool uses show up as counters on tokens.server_tool_use, for example {"web_search_requests" => 2}; unless the provider reports the charge, cost.total covers tokens only.

With acts_as_chat, every finished attempt is written to ruby_llm_usages before the assistant message callback runs, so a cancelled stream leaves an accounting row even when it leaves no message. Each row stores the operation, provider, model, status (succeeded, failed, or cancelled), the token buckets, and decimal costs. Costs are frozen when the attempt finishes, at the prices RubyLLM knew then, so run RubyLLM.models.refresh on a schedule; daily is plenty.

The table is good for reporting with one catch: SQL SUM skips NULL, so a sum of known costs looks like a complete bill even when it isn’t. Count the unpriced attempts next to it:

SELECT model,
       SUM(total_cost) AS known_cost,
       SUM(CASE WHEN total_cost IS NULL THEN 1 ELSE 0 END) AS unpriced_attempts
FROM ruby_llm_usages
WHERE created_at >= '2026-10-01'
GROUP BY model;

Every finished attempt also emits a usage.ruby_llm event with operation, provider, model, status, tokens, and cost, for embeddings, images, speech, transcription, moderation, OCR, and reranking as well as chat. Use it to feed metrics, billing, or finance systems.

The Tokens and Costs guide explains exactly what each total includes, and covers pricing usage yourself with cost_for.

Fallbacks, Cancellation, and Errors

Providers go down, and switching the model and redeploying is a slow way to respond. Declare backup models on the chat:

chat = RubyLLM.chat(model: "claude-sonnet-5")
  .with_fallbacks("gpt-5.6", "gemini-3.7-flash") # tried in order if Claude fails

response = chat.ask "Summarize this incident report."

If Claude fails with a rate limit, a server error, an overload, a timeout, or a dropped connection, RubyLLM retries the same request on GPT, then Gemini, with the same conversation, tools, schema, and settings. Once that generation is done, the chat goes back to Claude. Authentication and bad-request errors don’t fall back, because a wrong API key on provider A is not a reason to quietly bill provider B. on: replaces the list, and the defaults are a public constant, RubyLLM::Fallback::DEFAULT_ERRORS, so you can add RubyLLM::ContextLengthExceededError to hand long conversations to a model with a bigger window. Agents declare the same thing with fallbacks "gpt-5.6", "gemini-3.7-flash".

Fallbacks need credentials for each provider, and each fallback has to support what you’re asking for, such as structured output. before_fallback and after_fallback tell you when it happened. Streaming needs care, since the first model may have streamed half a sentence before it died. RubyLLM can’t take those chunks back, so the fallback starts a fresh assistant message, and fallback.chunks_yielded? tells you the user already saw something. Clear or mark the partial answer in your UI. The failed attempt is in the usage ledger, so chat.cost includes the half-answer you threw away.

chat.cancel stops a run at its next checkpoint by raising RubyLLM::CancelledError. The harder case is the stop button in a Rails app, where the stream runs in a job and the button lives in a different process:

class ChatsController < ApplicationController
  def cancel
    current_user.chats.find(params[:id]).cancel # stops the job in another process
    head :no_content
  end
end

class ChatStreamJob < ApplicationJob
  def perform(chat_id)
    chat = SupportAgent.find(chat_id)
    chat.complete do |chunk|
      chat.messages.last&.broadcast_append_chunk(chunk.content) if chunk.content
    end
  rescue RubyLLM::CancelledError
    # broadcast a "stopped" state if your UI shows one
  end
end

acts_as_chat writes the request to the chat’s cancelled column. The streaming job checks that column between chunks, at most once a second and outside the query cache, so it sees a write from another process without hammering your database. When it does, it clears the flag, raises, and removes the empty assistant row the stream created. The tokens already produced go into the ledger as a cancelled attempt. You don’t need a Redis key or a pub/sub channel. Cancellation is cooperative: a tool in the middle of arbitrary Ruby code finishes before the next checkpoint notices, and closing a browser tab doesn’t call cancel for you.

Provider errors come with subclasses like RateLimitError and PaymentRequiredError, all under RubyLLM::Error, and error.response keeps the HTTP response when there was one. Mistakes on your side (ConfigurationError, ModelNotFoundError, PendingToolCallsError) and CancelledError inherit straight from StandardError, so rescue RubyLLM::Error won’t treat a user pressing stop as a provider failure.

Rescue blocks around every call site get copied and drift apart, so agents declare their policy once with rescue_from, the way Rails controllers do: rescue_from RubyLLM::RateLimitError, RubyLLM::ServerError, with: :instrument_and_raise. The semantics are ActiveSupport::Rescuable’s, and handlers cover ask, complete, step, and the other verbs. Handlers wrap agent instances: SupportAgent.new.ask goes through them, but SupportAgent.find and SupportAgent.chat hand you the record or chat itself, which doesn’t. To get them on a persisted chat, use SupportAgent.new(chat: Chat.find(id), persist_instructions: false).

Thanks to @kieranklaassen for asking for fallbacks (#621), @sh1nj1 for cancellable streams (#607), and @skovy for rescue_from (#708).

The error handling guide covers model fallbacks and the error hierarchy, Rails streaming covers the stop button, and the agents guide covers handlers.

Prompt Caching

An agent sends the same system prompt, the same tool definitions, and the same 40-page contract on every turn. By turn twenty you’ve paid for that contract twenty times. Prompt caching lets the provider reuse a prefix it has already processed and bill the cache read at a fraction of normal input. RubyLLM 2.0 gives it one API:

chat = RubyLLM.chat(model: "claude-sonnet-5").with_caching # one switch for every provider
chat.with_instructions(review_guidelines)

response = chat.ask "Review this migration.", with: "db/migrate/20261001_add_billing.rb"
response.tokens.cache_write # first request: the prefix goes into the cache

Anthropic wants cache_control markers on content blocks. Bedrock Converse wants cache points. OpenAI-compatible APIs take a prompt_cache_key. OpenAI and Gemini also cache on their own. with_caching translates to each, so none of that ends up in your code. Options set the TTL (ttl: "1h"), the cache key, and the mode. Agents get the same thing with caching ttl: "1h", and with_caching(false) stops RubyLLM from sending cache controls, though it can’t turn off caching a provider does on its own.

Automatic caching guesses where your reusable prefix ends. When you know, mark it:

chat = RubyLLM.chat(model: "claude-sonnet-5").with_caching(ttl: "1h")

chat.with_instructions(analysis_prompt).cache_until_here # cache boundary
chat.add_message(role: :user, content: contract_text).cache_until_here

chat.ask "Does clause 14 conflict with the termination terms in clause 3?"
chat.ask "Which clauses mention liability caps?"

cache_until_here marks the latest message as a cache boundary, and RubyLLM sends your explicit boundaries plus automatic caching for the conversation growing after them. Anthropic, OpenRouter, Bedrock Converse, and selected OpenAI-compatible models understand boundaries; other providers keep their own caching behavior, so the same code runs everywhere. In Rails, the boundary is a cache_until_here column on the message, so it’s replayed every time the chat is loaded, from a controller or a job days later.

Gemini and Vertex AI also let you create a cache as a resource with a name and an expiry, and point any number of chats at it. RubyLLM.cache(content, model:, instructions:, ttl:) creates one, with_caching(id: cache) uses it, and RubyLLM::CachedContent.find(name, provider: :gemini) finds it again from another process, with renew and delete to manage it. One Gemini rule: a request that uses a cache can’t also send its own system instructions or tools, so put the instructions in the cache.

The providers still set the rules: minimum prefix length, how long the cache lives, which models support it. A hit is never guaranteed, and writing to a cache can cost more than plain input, so caching a prefix you never reuse costs more than not caching it. Check tokens.cache_read and cost.cache_write on your most repetitive workload. In Rails they’re recorded per attempt, so whether caching paid off is a query.

Thanks to @arunkumarry, whose issue and pull request for Anthropic and Bedrock caching got this started.

The prompt caching guide has the per-provider option table and the full cache lifecycle.

Batches

A user waiting for an answer needs an interactive request. An overnight job classifying ten thousand support tickets can wait. Anthropic, OpenAI, Gemini, Vertex AI, Bedrock, Azure, Mistral, xAI, OpenRouter, and Cohere all have batch APIs, and the major ones charge a lot less for it. Each has its own file format and steps for uploading, polling, and matching results. In RubyLLM 2.0, you write ordinary chats:

chats = tickets.map do |ticket|
  RubyLLM.chat(model: "claude-haiku-4-5")
    .with_instructions("Classify this ticket as billing, bug, or feature request.")
    .ask_later(ticket.body)
end

batch = RubyLLM.batch(chats) # submits them all at once
batch.id # save this

ask_later stages the question without sending anything, and RubyLLM.batch submits all of them in one go, with instructions, history, tools, and schemas. Later, from any process, RubyLLM::Batch.find(batch_id, provider: :anthropic) and batch.refresh.complete? tell you whether it’s done, and batch.messages comes back in submission order. A failed request is nil in messages, and batch.statuses tells you which slots succeeded, failed, or were cancelled, so you can resubmit the failures or finish them with chat.complete.

A batch generates one model turn. If the model asks for a tool, call run_tools on the chats and batch the ones that aren’t complete yet. That’s the agentic loop at batch prices.

When every input is a persisted chat, RubyLLM saves the batch in its own table, and a job can pick it up with nothing but the ID:

class BatchPollJob < ApplicationJob
  def perform(batch_id)
    batch = RubyLLM::Batch.find(batch_id) # Rails already knows the provider
    return self.class.set(wait: 10.minutes).perform_later(batch_id) unless batch.refresh.complete?

    batch.messages
  end
end

RubyLLM restores the provider and the chats, then saves each answer through the same callbacks as a normal ask, so broadcasts, after_create_commit hooks, and usage records run as if the user had been waiting. Run the job twice and you still get one answer per chat. Embeddings batch too: RubyLLM.embed_later stages them for a catalog backfill, and batch.results returns the vectors in submission order.

batch.cost is a RubyLLM::Cost. RubyLLM uses the provider’s reported cost when there is one, and otherwise prices each result at batch rates, which on Anthropic, OpenAI, Gemini, Bedrock, Azure, and Mistral means half the interactive price. The total stays nil until processing ends. On 2.0.0, batches from OpenAI reasoning models return a nil total because of how thinking tokens were priced; Marc Köhlbrugge fixed that in 2.1.

Use one provider per batch. Anthropic and xAI accept mixed models; the others want one model per submission. Bedrock batches take no tools or structured output, Cohere’s no structured output, and Bedrock and Vertex AI need a storage bucket for the batch files.

Thanks to @marckohlbrugge, @thomaswitt, @toddkummer, and @khasinski, whose requests and feedback shaped this.

The batch guide has the full restriction table and setup.

Owning the Transcript

A long conversation is full of things the user wants to keep and the model no longer needs to reread on every request. In 2.0 you can rewrite the history you send with chat.messages = messages_for_model. The setter takes Message objects, attribute hashes, or records that respond to to_llm, so you can summarize old turns, redact values, or drop a tangent, and the next request sends exactly what you put there. Your app knows what matters in a conversation better than a generic memory framework would.

This is practical because message.content is now a String or nil. Structured output is JSON text with a parsed reader, and files live on message.attachments. Conversations with tools need care: every tool call needs its result, and reasoning or provider-tool blocks have to stay with their message, so slicing the last four messages blindly can cut a call from its result, which providers reject.

On a Rails record, messages= is Active Record’s association writer. For a temporary rewrite, use chat_record.to_llm.messages = .... When the difference is permanent, for example the user sees everything and the model sees a redacted version, give RubyLLM its own association with acts_as_chat messages: :llm_messages, message_class: "LlmMessage". Your UI renders messages, and RubyLLM persists and sends llm_messages.

Some providers can compact a long conversation themselves. with_compaction(at: 100_000, instructions: "Keep every decision and every number.") sets the input-token trigger and steers the summary. Anthropic and OpenAI Responses write an opaque compacted block that RubyLLM keeps and replays; OpenRouter drops messages from the middle once the context is full, with no threshold and no summary. chat.compact compacts on demand through OpenAI, Azure, and xAI Responses, and chat.messages keeps every original message. Either way, the summarization is billed work, and RubyLLM counts it.

When you need to see or change the final payload, there’s before_request, and render shows the result without calling the model:

chat = RubyLLM.chat(model: "gpt-5.6")
  .ask_later("Summarize the changes.")

chat.before_request do |payload|
  payload[:metadata] = { review_id: "review-42" }
end

chat.render[:metadata] # => { review_id: "review-42" }
chat.complete

The hook runs after all of RubyLLM’s formatting and provider-option merging. It speaks the selected protocol’s wire format, so a hook written for OpenAI needs revisiting if you move to Anthropic. render also lets you test request shaping without network access or an API key. For a fixed field like this one, with_provider_options(metadata: ...) is simpler.

Finish reasons are now normalized symbols: :stop, :max_tokens, :tool_calls, or :content_filter, with stopped?, max_tokens?, tool_call_stop?, and content_filtered?. A truncated answer is :max_tokens whether the provider said length, max_tokens, or MAX_TOKENS, and Anthropic’s end_turn, Gemini’s STOP, and the Responses API’s completed are all :stop. Reasons RubyLLM doesn’t map, like Anthropic’s pause_turn, come through as symbols in the provider’s spelling. On Anthropic, :max_tokens also covers a full context window, so check with_max_output_tokens first and the transcript second.

Thanks to @mnort9 and @marksweston for pushing on transcript control, @fvaleye for compaction, and @trevorturk and @losingle for finish reasons.

Details are in the request control guide and the persistence guide.

Providers and Protocols

Cohere, Ollama Cloud, ElevenLabs, and Deepgram join the list, bringing the total to seventeen.

The bigger change is how much of each provider RubyLLM covers. I audited forty shared features across all seventeen providers, comparing 1.16 with 2.0. Built-in support went from 170 provider-feature pairs to 395, out of the 405 those providers offer. Those count provider-feature pairs, not distinct features or every endpoint: how much of each provider’s API you can reach from the same Ruby code. The coverage matrix shows every cell, with sources and the gaps that remain.

Under the hood, providers and protocols are now separate. A provider is the service you connect to; a protocol is the API it speaks. OpenAI defaults to the Responses API, and Vertex AI and Bedrock route each model to the API it actually speaks:

RubyLLM.chat(model: 'gpt-5.4')                              # OpenAI Responses API
RubyLLM.chat(model: 'gpt-5.4', protocol: :chat_completions) # same model, old API
RubyLLM.chat(model: 'claude-opus-4-6', provider: :vertexai) # Vertex AI, Anthropic protocol

Responses gives RubyLLM access to encrypted reasoning, provider-run tools, and newer OpenAI features. If you depend on Chat Completions, pass protocol: per chat or set config.openai_protocol = :chat_completions. Thinking has one API too: with_thinking uses the model’s defaults, or pass effort: or budget:, and read it back through response.thinking (guide).

A new provider can be a single small file. If yours is missing, generate a gem for it:

ruby_llm provider-gem Acme --api-base https://api.acme.ai/v1

I covered the design in Providers, Protocols, and Provider Gems.

Rails Table Ownership

A fresh 1.x install put four models in your app: Chat, Message, Model, and ToolCall. In 2.0 it’s two:

# app/models/chat.rb
class Chat < ApplicationRecord
  acts_as_chat
  belongs_to :user
end

# app/models/message.rb
class Message < ApplicationRecord
  acts_as_message
  has_many_attached :attachments
end

Chats and messages are your product’s conversations, so they stay yours: users, scopes, authorization, titles, retention. Everything else lives in tables RubyLLM owns under a ruby_llm_ prefix: ruby_llm_models for the registry, ruby_llm_tool_calls for tool requests, approval decisions, and links to their results, ruby_llm_usages with one row per provider attempt, and ruby_llm_batches. You read them through the same API as in plain Ruby: RubyLLM.models, message.tool_calls, message.tokens, chat.cost, and RubyLLM::Batch.find.

In 1.x the generator copied Model and ToolCall into your app, and from then on they were your problem. When RubyLLM needed to store something new about a tool call, you had to update a class you didn’t write and never called directly. 2.0 stores a lot more: approval decisions, a usage row per attempt, batch state. Shipping that as “please update these four files” would have made a painful upgrade, and the next feature would have needed another one.

Rails has a precedent. You don’t have an ActiveStorageBlob in app/models; Active Storage owns its records and you use them through has_many_attached. RubyLLM’s tables are ordinary tables created by ordinary migrations, not an engine. The record classes behind them are internal and the readers are public, so RubyLLM can change its storage without your app noticing. Every feature in this post that needed new columns got them without asking you to maintain another model.

Usage rows also make your own reporting plain Active Record: has_many :ruby_llm_usages, through: :chats on User, then current_user.ruby_llm_usages.sum(:total_cost). Keep the nil caveat from cost and usage in mind.

If you put your own columns on the old models table, like availability flags, default models, or admin pricing, those are product features. Give them a table of your own keyed by provider and model ID. The upgrade keeps the columns, but RubyLLM won’t maintain them.

The Rails persistence guide shows every association and reader.

Prompt Templates

A two-page prompt inside a Ruby heredoc is hard to work with. The prompt changes often, the class around it rarely does, and every diff is paragraphs of English between def and end. Prompts are templates, so in 2.0 they get a directory the way views do, app/prompts, and RubyLLM.render_prompt renders them:

<%# app/prompts/support/instructions.txt.erb %>
You are a support assistant for <%= product_name %>.
The current customer is <%= customer_name %>.
Answer with concise, practical steps.
instructions = RubyLLM.render_prompt(
  "support/instructions",
  product_name: "BillingHub",
  customer_name: current_user.name
)

It reads an ERB file, renders it with your locals, and returns a String. It doesn’t call a model or create a chat, so use the string as instructions, as a user message, or as input to RubyLLM.embed. Lookup is relative to Rails.root in Rails and the current directory in plain Ruby. A missing local makes ERB raise, and a missing file raises RubyLLM::PromptNotFoundError, so you can test every prompt without an API key. Prompt files are code, since ERB runs Ruby: keep prompt names in your code and pass user input only as locals.

Agents have used app/prompts since 1.12, and now share the same renderer. In 2.0 a named agent picks up its prompt by convention: WorkAssistant reads app/prompts/work_assistant/instructions.txt.erb if it exists, and starts without instructions if it doesn’t. When the prompt is required, declare instructions { prompt("instructions") } and a missing file raises. Locals can be lambdas that run with access to the chat and the agent’s inputs, and instructions can stack: instructions append: true, persist: false adds text, like today’s date, that stays out of the saved transcript.

A Rails-backed agent saves its instructions when create! makes the chat, and WorkAssistant.find(id) applies the current configuration without rewriting what’s saved. When old conversations should pick up new wording, WorkAssistant.sync_instructions(chat) re-renders and saves them (it was sync_instructions! in 1.x). A bare instructions call used to require the conventional prompt; in 2.0 it’s only the reader.

Rails engines can ship prompts. RubyLLM::Prompt.roots is an ordered list of prompt directories, like Action View’s view paths, with the app’s app/prompts always first. An engine appends its own directory, and the host app overrides any engine prompt by creating the same path in app/prompts. This came from @adrianthedev. Thanks to @kryzhovnik, who extracted the renderer out of Agent’s private methods, which made this possible.

The prompt rendering guide has the details, and the agents guide covers the class-based conventions.

Workflows and Instrumentation

A research agent calls a model, searches the web a few times, and hands its notes to a writing agent. The Ruby is short, but the logs show dozens of separate events with nothing tying them to the article they produced. 2.0 lets you name the work:

RubyLLM.workflow("Write article", id: "article-42") do |workflow| # tags every event inside
  notes = workflow.step("Research") do
    ResearchAgent.new.ask(topic).content
  end

  workflow.step("Draft") do
    WriterAgent.new.ask(notes).content
  end
end

Every RubyLLM event inside the block carries workflow_id and workflow_name, and inside a step, workflow_step_id and workflow_step_name. Model calls, tool calls, and usage rows for retries are all tagged. The blocks return their normal values, so notes is a String.

Many AI frameworks come with a graph DSL: nodes, edges, a state object, and a runtime that executes it. In Ruby, a sequence is method calls, a branch is a case, and a retry is retry. What’s missing is a way to see that control flow afterwards, and that’s all RubyLLM.workflow does. It doesn’t persist progress, schedule anything, or retry. For durability, use the loop’s verbs and your job queue.

1.16 introduced instrumentation with five events. 2.0 has twenty: every model operation, per-attempt usage, and the workflow and step wrappers. In Rails they go through ActiveSupport::Notifications, so you subscribe the way you’d subscribe to sql.active_record. Subscribe to usage.ruby_llm and group by workflow_id to get the cost of a piece of work, failed attempts included. Outside Rails, set config.instrumenter to anything that responds to instrument(name, payload) and yields. If you wrote subscribers for 1.16, tokens moved under payload[:tokens].

One-shot operations take metadata:, which lands in payload[:metadata] and is never sent to the provider. Workflows take it too, and nested events get it as workflow_metadata. Steps and workflows nest, recording parent IDs, so a subscriber can rebuild the execution tree of a run. Context lasts exactly as long as the block, so it doesn’t follow your work into a job that runs tomorrow; save the ID and open the workflow again with the same id:. When you run work concurrently yourself, open the step inside each task.

Payloads include message content, tool arguments, and provider responses, so export those only where your policy allows it. The instrumentation guide lists every event and payload field.

Coding Assistant Skill

The gem ships a RubyLLM skill for coding assistants, maintained alongside the code and guides. Install the copy that matches your bundle:

npx skills add "$(bundle show ruby_llm)" --skill rubyllm

Your assistant stops writing 1.x method names from its training data and stops rebuilding things RubyLLM already does. Setup details.

Upgrading

There are two jobs: update your Ruby code, and, if you use Rails persistence, migrate your stored records. Plain Ruby apps only do the first. Your chats and messages keep their IDs and relationships, nothing is deleted until you run cleanup, and every migration phase can be retried.

Start on 1.16. If you’re on something older, step through the minor releases first, because the generator expects the schema 1.16 produced. While still on 1.16, set config.deprecation_behavior = :raise in your test environment and fix whatever breaks, and if you still have config.use_new_acts_as = false, switch to the association-based acts_as API, the only one 2.0 has.

Pin the 2.0 series. Then run bundle update ruby_llm:

gem "ruby_llm", "~> 2.0.0"

The 1.16 upgrade generator, its migration helpers, and the copy-mode tasks ship with 2.0 only. From 2.1 on, each release carries just the upgrade from the release before it. Finish this one, cleanup included, before you move to 2.1.

Update your calls. Most changes are renames, because each concept now has one name: with_tool becomes with_tools; tool desc, param, and params become description, parameter, and parameters; with_params becomes with_provider_options; response.input_tokens becomes response.tokens.input; on_new_message and on_tool_call become before_message and before_tool_call; create_user_message becomes ask_later; RubyLLM::Schema becomes Schematist::Schema. Rails controller params have nothing to do with the params: rename, so change RubyLLM call sites one by one rather than with search-and-replace. If you tried a release candidate, rename with_server_tools to with_provider_tools.

A few behavior changes matter more than names. Finish reasons are Symbols. Switches like with_thinking, with_caching, with_citations, and with_compaction take false to turn off, and value setters like with_temperature take nil to reset. The new before_ and after_ callbacks stack, where the old on_* ones replaced each other. OpenAI uses the Responses API, so Chat Completions-only options like response_format need the Responses shape or config.openai_protocol = :chat_completions. The temperature you set is the temperature sent; 1.x sometimes rewrote it for certain models.

Choose rename or copy. Rename, the default, moves your existing model and tool-call tables into RubyLLM’s ownership. It’s fast, needs AI activity paused for all three phases, and going back means restoring your backup with the matching app build. Copy (--mode copy) builds the 2.0 tables next to the 1.16 ones, runs preparation and backfill while 1.16 keeps serving traffic, and keeps a supported route back to 1.16 with ruby_llm:upgrade:rollback and ruby_llm:upgrade:resume. On 100,000 synthetic chats with a million messages on PostgreSQL:

Mode Total migration time AI downtime
Rename 20 s 20 s
Copy 136 s 4 s

Those are medians from my migration benchmark, not a promise about your database. Copy mode generates a compatibility concern and initializer that must run in both builds, and its guards don’t cover update_columns, bulk SQL, direct deletes, or attachment purges. Its migrations also need a direct database connection or a session-mode pool. If you don’t need the way back or the shorter pause, use rename.

Generate and rehearse. bin/rails generate ruby_llm:upgrade (add --mode copy and any custom class mappings) writes three migrations: prepare, backfill, and finish. Backfill works in batches of 10,000 messages and saves its progress, so a rerun skips finished batches and never counts historical usage twice. Rehearse on a recent copy of production with the exact build you’ll deploy, time it, and check message counts, tool-result links, and token counts. Take a backup you’ve actually restored before.

Run it. For rename, finish or cancel conversations waiting on tool results, stop everything that touches AI, deploy 2.0, and run bin/rails db:migrate and bin/rails ruby_llm:load_models. Copy mode deploys its compatibility code to 1.16 first, runs prepare and backfill from the 2.0 build while 1.16 serves, then pauses only for finish. If your deploy’s migration step has a timeout, generate the phases one at a time with --phase.

Delete two models. Remove Model and ToolCall with their acts_as_model and acts_as_tool_call declarations, drop model: and tool_calls: from the remaining macros, and remove config.model_registry_class and config.use_new_acts_as. Don’t drop the old tables yourself. The backfill creates one usage row per historical response with recorded usage, but it can’t recover retries or prices 1.16 never recorded.

Clean up later. Once you trust 2.0 in production, generate --phase cleanup in a later release. In copy mode, run bin/rails ruby_llm:upgrade:finalize first, which closes the way back. Then delete the 2.0 upgrade migrations from db/migrate.

The migrations have no down, and changing the gem version back does not roll back the database. The complete 2.0 upgrade guide has every rename, the copy-mode procedure step by step, incomplete tool calls, and custom class names.

Get It

bundle add ruby_llm

What’s New in 2.0 has working examples of everything above, the upgrade guide walks a Rails app from 1.16, and the documentation covers the rest.

Thanks to everyone who tested the release candidates, filed issues, and sent code, including 22 people who made their first contribution during this cycle.

RubyLLM 2.1 comes out on October 8, at Deccan Queen on Rails in Pune.

Newsletter