Welcome
Welcome to the homepage for llm.rb
llm.rb is an advanced runtime for building capable AI applications on CRuby. By default it has zero runtime dependencies although certain functionality – such as ActiveRecord support – require optional dependencies that are opt-in.
Features
The runtime supports OpenAI, OpenAI-compatible endpoints, Anthropic, Google Gemini, Mistral, DeepSeek, DeepInfra, xAI, Z.ai, AWS Bedrock, Ollama, and llama.cpp. It has first-class support for streaming, tool calls, MCP and A2A, embeddings, vector stores and the RAG pattern.
There are multiple HTTP backends to choose from, tools can be run concurrently or in parallel via threads, async tasks, fibers, ractors, and fork, and it is also possible to make a tool call while the model is still streaming.
The runtime builds on top of three core concepts: providers, contexts, and agents, so once you learn the fundamentals, everything else falls into place naturally. And once you learn llm.rb, you will also be able to use mruby-llm and wasm-llm because the API is pretty much identical.
Install
Source code and releases are available from github.com/r-uby-dev/llm.
$ gem install llm.rb
Quick start
LLM::Agent
The LLM::Agent class is the default high-level interface,
and it is recommended for most use-cases. It manages tool execution
automatically, guards against infinite loops, manages conversation
state, and much more.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, stream: $stdout)
agent.talk "Hello world"
LLM::Context
The LLM::Context class is at the heart of the runtime
and it is what LLM::Agent uses under the hood.
It requires that the tool call loop be managed manually –
sometimes that can be useful, but usually for advanced use-cases.
If you're new to llm.rb, try LLM::Agent first.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
ctx = LLM::Context.new(llm, stream: $stdout)
ctx.talk "Hello world"
LLM::REPL
REPL: Agent
The LLM::Agent#repl method allows
an agent to spawn a read-eval-print loop that can be useful while
developing or operating agents. It can be used to debug tool calls,
confirm an agent has done what was expected, or improve an agent by
asking questions about what it has done up to that point.
This feature requires that the curses and kramdown libraries are installed and available to require.
The TUI displays a status line with a context-usage bar and cost counter, a scrollable transcript with markdown rendering, and a multi-line input area. The UI stays responsive while the model is generating a response.
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm)
agent.repl
REPL: State
The path: option lets you resume a session across REPL runs by
reading and writing the agent state from a file.
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm)
agent.repl(path: "session.json")
REPL: Tools
The tools option lets you attach additional tools for the
duration of the session. This is in addition to any tools that might
already be associated with an agent.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm)
agent.repl(tools: [Debugger])
A number of optional tools are distributed as part of llm.rb. They
power the agents in the agents/ directory, so they're optimized
for developer tasks. The following example starts a read-eval-print
loop with all builtin tools available.
require "llm"
require "lll/tools"
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm)
agent.repl(tools: LLM::Tool.subclasses)
REPL: Skills
The skills option lets you load extra skill directories
without attaching them to an agent permanently.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm)
agent.repl(skills: [__dir__])
REPL: Tracer
By default the tracer is disabled for the session. The tracer
option can re-enable it, and setting it to true uses the tracer
already associated with the agent.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
tracer = LLM.logger(llm, path: "agent.log")
agent = LLM::Agent.new(llm, tracer:)
agent.repl(tracer: true, tools: [Debugger])
REPL: Commands
Commands are recognized by a / prefix and are backed by
LLM::Repl::Command,
which can be subclassed to add custom commands. Once you create a
subclass, it is automatically added to the repl. A command can have
zero or more parameters, and all parameters are presumed to be a
String, at least for now.
require "llm"
require "llm/repl"
class Greeter < LLM::Command
name "greet"
description "Greets the given name"
parameter :name, String, "The person's name"
required %i[name]
def call(name:)
write("Welcome #{name}!\n")
end
end
REPL: Input
The input area supports keyboard shortcuts, paste mode for multi-line
prompts, and slash commands such as /exit.
LLM::Tool
Subclasses of LLM::Tool are plain Ruby classes with
an optional set of typed parameters.
The model can choose to
call them on your behalf, and they're one of the most powerful features
for extending the feature set or abilities of a model.
class ReadFile < LLM::Tool
name "read-file"
description "Read a file"
parameter :path, String, "The filename or path"
required %i[path]
def call(path:)
{contents: File.read(path)}
end
end
LLM::Stream
Streams can be simple IO objects or subclasses of
LLM::Stream with structured callbacks for content,
reasoning, tool calls, tool returns, and compaction.
class MyStream < LLM::Stream
def on_content(content)
print content
end
def on_reasoning_content(content)
warn content
end
end
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, stream: MyStream.new)
agent.talk "Explain Ruby fibers."
LLM::MCP
The Model Context Protocol (MCP) has first-class support
in llm.rb. The stdio and http transports work out of the
box. MCP tools are translated into subclasses of
LLM::Tool that can be used with LLM::Context
or LLM::Agent.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
mcp = LLM::MCP.stdio(argv: ["ruby", "server.rb"])
agent = LLM::Agent.new(llm, stream: $stdout, tools: mcp.tools)
agent.talk "Run the tool"
LLM::A2A
The Agent 2 Agent (A2A) protocol has first-class support
in llm.rb. The http and jsonrpc transports work out of the
box. A2A skills are translated into subclasses of
LLM::Tool that can be used with LLM::Context
or LLM::Agent.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
a2a = LLM::A2A.rest(url: "https://remote-agent.example.com")
agent = LLM::Agent.new(llm, stream: $stdout, tools: a2a.skills)
agent.talk "Run the skill"
RAG
Most providers offer an embedding model that can be used for semantic search, or similarity search. An embedding model can generate embeddings that can then be stored in a database that is optimized for storing and querying vectors, such as SQLite's sqlite-vec or PostgreSQL's pg-vector.
llm.rb also includes support for OpenAI's vector store API. It provides a vector database as a HTTP service but we won't cover that here.
require "llm"
llm = LLM.openai(key: ENV["KEY"])
body = "llm.rb is Ruby's capable AI runtime."
embedding = llm.embed([body]).embeddings.first
Document.create!(
title: "llm.rb",
body:,
embedding:,
)
Concurrency
The runtime supports five different concurrency strategies that have different attributes. The choice between all of them often depends on the requirements of your application.
IO-bound tools are a good fit for the :task, :thread,
and :fiber strategies while true parallelism can be achieved
with the :fork and :ractor strategies. The
:fork strategy also provides a separate process that offers
isolation from its parent.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
tools = [FetchNews, FetchStocks, FetchFeeds]
agent = LLM::Agent.new(llm, tools:, concurrency: :fork)
agent.talk "Run the tools in parallel"
ORM
Because both LLM::Context, and LLM::Agent
can be serialized to JSON and stored in a simple string both ActiveRecord
and Sequel support can be implemented within a single column on a single row.
The runtime
includes first-class support for both ActiveRecord and Sequel, and
for both Rack-based applications and Rails-based applications. On
databases where it is supported, such as PostgresSQL, the column can be
optimized by using the jsonb type.
require "active_record"
require "llm"
require "llm/active_record"
class Agent < ApplicationRecord
acts_as_agent do |agent|
agent.model "deepseek-v4-pro"
agent.instructions "solve the user's query"
agent.tools [Research, FinalizeResearch, ActOnResearch]
end
private
##
# By convention, this method defines the provider
# for a model. If neccessary, it can be renamed and
# configured via `provider: :your_method` instead.
def set_provider
LLM.deepseek(key: ENV["KEY"])
end
##
# By convention, this method should return what is
# given as the second argument to `LLM::Context` or
# `LLM::Agent`.
#
# Often, there is no need to set it, so it can be left
# undefined or it can be reassigned in the same way as
# `set_provider`. For example: `context: :your_method`
def set_context
{}
end
end
agent = Agent.create!
agent.talk "perform research"
Learn more
The Quick Start guide is intended to be a quick overview of
the runtime, it lacks depth, and omits certain features. The
deepdive provides a thorough,
deep overview of the runtime that is recommended as the
next stop for learning more about
llm.rb.
FAQ
What providers does llm.rb support?
Cloud
The following cloud-based providers are available to choose from.
In no particular order:
πΊπΈ OpenAI
πΊπΈ DeepInfra
πΊπΈ xAI
πΊπΈ Google (Gemini)
πΊπΈ AWS bedrock
πΊπΈ Anthropic
π¨π³ DeepSeek
π¨π³ zAI
πͺπΊ Mistral
Weights
The following providers provide access to open-weight models.
In no particular order:
πΊπΈ DeepInfra
πΊπΈ AWS bedrock
π¨π³ DeepSeek
π¨π³ zAI
πͺπΊ Mistral
Local
The following providers can be run locally on your own hardware.
In no particular order:
- Ollama
- Llamacpp
I have a limited budget. What should I do?
There are a few options. The first option is to host your own model and use the Ollama or llama.cpp providers. This can be difficult though because a capable model requires hardware that can match it. If you have the ability to self-host, this would be my first option.
The second option is DeepSeek. The deepseek-v4-flash model costs pennies to use. And llm.rb has been optimized for DeepSeek. For example, DeepSeek does not have image generation capabilities but on the llm.rb runtime it does, at least for vector graphics.
The same is true for structured outputs. DeepSeek does not support them in the same way as OpenAI or Google, but the llm.rb runtime makes it appear as though it does through the json_object response type.
If you are on a budget, DeepSeek is hard to beat.
Can I download llm.rb via a decentralized network?
You can.
We are on the radicle.network.
Every commit that lands on GitHub also lands on Radicle.
Our repository ID is z2PtfQ6dYwyYaW2aGrztG1sMyDmCE.
Browse it on the web.
Applications
mruby-llm is a port of llm.rb to the mruby runtime, and it
has been used to build novel applications that are available to the general public
via SSH.
| Application | Try it | Runtime |
|---|---|---|
| matz | ssh matz@r.uby.dev |
mruby-llm |
| robert | ssh robert@4.4bsd.dev |
mruby-llm |