RubyLLM 2.1: MCP, Judgments, Evaluations, and Less Work on Every Call

RubyLLM 2.1: MCP, Judgments, Evaluations, and Less Work on Every Call

RubyLLM 2.1 is out. I released it on stage at Deccan Queen on Rails in Pune.

It builds on 2.0 and adds an MCP client, typed judgments, evaluations, tool progress, OpenTelemetry tracing, and a nineteenth provider. It’s also faster, with no code changes. In this post:

An MCP Client Where You Decide What the Model Sees

In RubyLLM 2.1 an MCP server is a Ruby class you own: it lives in your repo, and it says exactly which tools the model gets.

class Linear < RubyLLM::MCP
  url "https://mcp.linear.app/mcp"
  inputs :user
  oauth owner: :user

  only :list_issues, :get_issue, :create_issue # the model sees these three
  requires_approval :create_issue
end

chat = RubyLLM.chat.with_mcp(Linear.new(user: current_user)) # one connection per user
chat.ask "What's blocking the release?"

To try a server before shaping it, connect from the console. Every tool becomes a Ruby method:

>> docs = RubyLLM.mcp(url: "https://learn.microsoft.com/api/mcp")
>> docs.microsoft_docs_search(query: "Azure Blob Storage").text
=> "..."

The same class can rename tools, rewrite their descriptions, fix arguments, and wrap results, which is how 2.1 follows the advice from April that your agent’s context window is not a junk drawer: prototype with MCP, then craft the tools you control. Before this, ruby_llm-mcp by @patvice carried MCP for RubyLLM users for a long time. Thank you. If you use it, remove it before updating, since it defines the same RubyLLM::MCP constant. The MCP Client guide covers OAuth, input requests, resources, prompts, MCP Apps, and Tasks.

Judgments

Judgments answer questions about your own data (is this urgent, which team owns it) with probabilities a model measured, instead of a confidence number it wrote as text:

class TicketTriage < RubyLLM::Judge
  probability :urgent, "Does this need attention today?"

  choice :department, "Which team should handle this?" do
    billing   "Payments and refunds"
    technical "Bugs and integrations"
    other     "Everything else"
  end

  score :frustration, "How frustrated is the customer?",
    ["Calm", "Frustrated", "Angry"]
end

judgment = TicketTriage.judge("Please refund the duplicate charge today.")
judgment.urgent.probability # => 0.96
judgment.department.choice  # => :billing

probability, choice, and score questions are all asked over the same input in one request, and your code acts on them with thresholds you pick. Judges default to Jev models from TypeSafe, a new built-in provider. Kieran Klaassen added OpenAI’s gpt-6-luna through the Decisions API (#1008). See the Judgments guide.

Evaluations

An evaluation runs your agent on cases with known answers and has a model grade each response, so you can tell whether a prompt or model change made things better. Cases live in a YAML file next to the class, and bin/rails "ruby_llm:eval[DocsEvaluation]" runs them and exits non-zero on failure, so it works in CI:

class DocsEvaluation < RubyLLM::Evaluation
  evaluation :correctness # keeps the default check
  evaluation :grounded, "Every claim is supported by the documents in metadata"
  evaluation :cites_sources, "The answer links to at least one document"

  def perform(question)
    DocsAgent.new.ask(question)
  end
end

Evaluations can also grade tool calls with plain Ruby assertions, use an Agent or a Judge as the grader, and run as RSpec or Minitest tests. See the Evaluations guide.

Faster by Default

2.1 does less work on every call, with no code changes. RubyLLM’s own work, 2.0.0 against 2.1 on the same machine:

Workload 2.0 2.1
Stream a 2 MB event that arrives in 16 KB pieces 116 ms 2.0 ms
Stream 500 Perplexity chunks that each cite 20 sources 137 ms 17 ms
Ask a Bedrock chat with 200 messages of history 2.9 ms 0.40 ms
Memory kept by a streamed 40-turn chat with a 256 KB image 28 MB 0.48 MB
Eight threads loading the model registry at once 273 ms 32 ms

Connections are now shared across calls, threads, and fibers. To keep them open, pick a persistent adapter (from the faraday-net_http_persistent gem):

config.faraday_adapter = :net_http_persistent # keeps connections open

The benchmarks need no API keys: clone RubyLLM and run bundle exec rake "benchmark:compare[v2.0.0]" to compare any version with your checkout. The connection guide covers adapters.

Tools That Report Progress

A slow tool can now tell your user what it’s doing instead of leaving them with a spinner:

class ReadReport < RubyLLM::Tool
  def execute(url:)
    pages = Scanner.pages(url)
    pages.each_with_index.map do |page, index|
      progress "Reading page #{index + 1} of #{pages.size}", value: index + 1, total: pages.size # report progress
      page.text
    end.join("\n")
  end
end

chat.with_tools(ReadReport).after_tool_progress do |tool_call, progress| # receives every update
  puts "#{tool_call.name}: #{progress.message}"
end

MCP tools report through the same callback. See Reporting Progress.

OpenTelemetry Tracing

One line sends every model call, tool run, and workflow as a span to the tracing backend you already use:

RubyLLM::OpenTelemetry.enable

Spans join the current trace, so the HTTP call your tool makes nests under that tool. They carry metadata such as models, tokens, and tool names, and never prompts, responses, or tool arguments. The OpenTelemetry guide has the full span reference.

Smaller Changes

  • Hetzner is provider number nineteen: RubyLLM.chat(model: "Qwen3.8-27B", provider: :hetzner).
  • Per-tenant agents: a context block can use the agent’s inputs, so each workspace can bring its own API key. Thanks to mikemikimike (#903).
  • A secondary database can hold RubyLLM’s supporting records next to your chats.
  • Usage beyond chats: in Rails, embeddings, transcriptions, and other one-shot calls go into the usage ledger too.
  • error.request_shape lists every turn’s parts and sizes when a provider rejects a request.
  • Provider uploads are reused across processes in Rails, so a large PDF isn’t uploaded again on every job.
  • Perplexity chat runs on the Agent API, since Sonar retires.
  • Ruby 3.2 or later is required, since 3.1 reached end of life.
  • JSON 3 is allowed, thanks to Filipe Kalicki (#968), who also made model lookups use an index (#981).

Upgrading from 2.0

gem "ruby_llm", "~> 2.1.0"
bundle update ruby_llm
bin/rails generate ruby_llm:upgrade
bin/rails db:encryption:init   # only for MCP OAuth, if your app has no encryption keys yet
bin/rails db:migrate

Most 2.0 apps need no code changes. The few that do are listed in the upgrade guide, and everything new is in What’s New in 2.1.

Thanks to everyone who sent code for this release: Andrii Furmanets, Andrey Samsonov, Andy Wang, Anton Kopylov, Filipe Kalicki, Guilherme Lages Santos, Islam Gagiev, Kieran Klaassen, Marc Köhlbrugge, mikemikimike, Mikhail Topolskiy, Muhammad Zain Ul Abidin, Paul Arterburn, and Viktor Schmidt.

Newsletter