Repository files navigation

RubySage

Gem VersionCILicense: MIT

Chat with your Rails codebase. Drop in a helper, scan your app, ask questions, get answers grounded in your actual code with citations.

Why

You've been on a Rails team for six months and nobody can explain how the subscription billing flow works because Steve left in March. You ask the codebase. RubySage answers — with file citations.

How it works

  1. Scan: a daily rake task walks your codebase and produces per-file artifacts — small structured records with summaries, public symbols, and route mappings. Stored in your app's database. Secret values redacted.
  2. Retrieve: when someone asks a question, RubySage retrieves the most relevant artifacts (keyword + symbol matching, page-context boosted).
  3. Answer: the relevant artifacts go into a prompt with the question. The LLM answers, citing the specific files and classes it used.

No giant snapshot stuffed into every prompt. No invented file paths. Just retrieval grounded in your actual code.

Install

Add to your Gemfile:

gem"ruby_sage"

Then:

bundle install
rails generate ruby_sage:install
rails db:migrate

Edit config/initializers/ruby_sage.rb to set your provider and API key.

Quickstart

# config/initializers/ruby_sage.rbRubySage.configuredo |config|
config.provider=:anthropicconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.auth_check=->(c){c.current_user&.admin?}end

Run your first scan:

bundle exec rake ruby_sage:scan

Drop the widget into your layout:

<%# app/views/layouts/application.html.erb %><body><%=yield%><%=ruby_sage_widget%></body>

Open your app, click the floating button, ask "what does the PostsController do?". Done.

Providers

Anthropic (recommended)

Uses prompt caching from day 1. The retrieved-artifact context block carries a cache_control: ephemeral marker — subsequent questions within ~5 minutes hit the cache and cost roughly 10× less.

config.provider=:anthropicconfig.model="claude-sonnet-4-6"config.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)

OpenAI

Works fine. No prompt caching in V1.

config.provider=:openaiconfig.model="gpt-4.1"config.api_key=ENV.fetch("OPENAI_API_KEY",nil)

Running scans

In development

Run manually after material code changes:

bundle exec rake ruby_sage:scan

Files unchanged since the previous scan are skipped via digest cache, so re-scans are cheap.

Use your local coding agent (no API key needed)

If you already use Claude Code, Codex, or Cursor, your agent can produce the summaries instead of RubySage paying for them. Two-step flow:

bundle exec rake ruby_sage:scan:plan

Writes tmp/ruby_sage/manifest.json (one entry per scanned file, with redacted contents inlined) and tmp/ruby_sage/INSTRUCTIONS.md. Tell your agent:

Read tmp/ruby_sage/INSTRUCTIONS.md and follow it.

The agent writes tmp/ruby_sage/summaries.json. Then:

bundle exec rake ruby_sage:scan:apply

A new completed Scan lands in your DB. Files unchanged since a prior scan reuse their cached summary, so the agent only summarizes what actually changed. Your ANTHROPIC_API_KEY / OPENAI_API_KEY is never used here — the cost lives in your existing agent subscription.

In production

Daily cron is the typical pattern:

# config/schedule.rb (using the `whenever` gem) — or your Heroku Scheduler / GitHub Actions cronevery:day,at: "4am"dorake"ruby_sage:scan"end

Pre-bake in CI, ship to prod

Expensive scans don't have to run in production. Bake them in CI:

# .github/workflows/snapshot.yml
- run: bundle exec rake ruby_sage:scan
- run: bundle exec rake ruby_sage:export_artifacts > artifacts.json
- uses: actions/upload-artifact@v4with:
name: ruby_sage_snapshotpath: artifacts.json

Then on deploy:

bundle exec rake ruby_sage:import_artifacts < artifacts.json

Production skips LLM summarization entirely — zero token spend on scans.

The agent-driven flow ships to production the same way. Run scan:plan + agent loop locally or in CI, then check the resulting manifest.json + summaries.json into a deploy artifact and run scan:apply on prod with MANIFEST=... and SUMMARIES=... pointing at the shipped files.

Cost

Rough numbers, USD, with Anthropic at current Sonnet pricing. Your mileage will vary by codebase size and question patterns.

OperationFrequencyCost
Initial full scan (200-file Rails app)once~$0.10–$0.50
Daily incremental scan (50 changed files)per day~$0.05–$0.20
User question (with prompt-cache hit)per question~$0.01–$0.05
User question (cache miss / first of session)per question~$0.10–$0.30

Pre-baking in CI moves the daily-scan cost off your production token budget entirely.

AI agent integration

RubySage's retrieval layer is a public Ruby API, not just a widget. Wire it into Cursor, Claude Code, Codex, or your own coding agent to give the model indexed knowledge instead of dumping the whole codebase into context.

result=RubySage.context_for("how does subscription billing work?")result[:artifacts]# => relevant Artifact recordsresult[:citations]# => [{path:, kind:, score:, snippet:}, ...]result[:scan_id]# => Scan id used

Or hit the JSON endpoint:

curl -X POST https://your-app.com/ruby_sage/internal/retrieve \
-H "Content-Type: application/json" \
-d '{"query":"subscription billing"}'

A typical Cursor / Claude Code session can spend 50–200K input tokens orienting the model to a Rails codebase before any real work happens. Swap that for a 3K-token retrieval call and your dev token bill drops by an order of magnitude.

Generate onboarding docs

bundle exec rake ruby_sage:onboard

Writes two files based on the latest scan:

  • docs/ONBOARDING.md — for human developers joining the team. Tech stack, data model, key workflows, where to start reading, gotchas.
  • docs/AGENT_PRIMER.md — for AI coding agents. App in one sentence, stack, domain model, service layer, patterns to follow / what NOT to do. Kept under 600 words so it fits any agent's context prefix.

A new developer or AI agent runs this once and has structured context in under a minute.

CLI chat

bundle exec rake "ruby_sage:ask[how does authentication work?]"# or
QUERY="how does authentication work?" bundle exec rake ruby_sage:ask

Prints the answer + source file citations to STDOUT. Useful for shell-based AI agents that need a quick lookup without opening a browser.

Admin dashboard

Visit /ruby_sage/admin/scans for scan history, artifact counts by kind, and a "Scan now" button. Visit /ruby_sage/admin/artifacts to browse the indexed file list with summaries and public symbols. Both go through the same auth gate as the chat endpoint.

Configuration reference

RubySage.configuredo |config|
# Providerconfig.provider=:anthropic# :anthropic | :openaiconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.summarization_model="claude-haiku-4-5"# Authorization (server-side, enforced before_action)config.scope=:admin# :admin | :signed_in | :public_rate_limitedconfig.auth_check=->(c){c.current_user&.admin?}# Mode (shapes prompt + filters artifacts by audience)config.mode=:developer# :developer | :admin | :user# Audience scopingconfig.user_facing_paths=["app/views/help/**/*"]# additively tag :userconfig.audience_for=->(attrs){ ... }# full override# Admin database queries (read-only SELECT via tool loop)config.enable_database_queries=falseconfig.query_scope=->(c){"organization_id = #{c.current_user.organization_id}"}config.query_connection=->(_c){ReadOnlyDatabase.connection}config.max_query_rows=100config.query_timeout_ms=5_000config.tool_loop_max_iterations=5# Optional CSP nonce hookconfig.csp_nonce=->(c){c.content_security_policy_nonce}# Scannerconfig.scanner_include=[...]# see lib/ruby_sage/configuration.rb for defaultsconfig.scanner_exclude=[...]config.scan_retention=7# HTTPconfig.request_timeout=30config.max_retries=2end

Security

  • Server-side auth. Every chat / retrieve request goes through before_action :authorize_ruby_sage!. The widget UI is just UX; the endpoint is the gate.
  • Secret redaction. YAML values for keys matching (api_key|secret|password|token|access_key|private_key|client_secret) are replaced with [REDACTED] at scan time. ENV[...] symbol references are preserved (the model needs to know which dependencies exist), but values never are. credentials.yml.enc is excluded entirely.
  • Provider data policies. Anthropic and OpenAI receive your code summaries (not raw source by default) when you ask questions. Read each provider's data retention policies before scanning sensitive code.
  • Supported Ruby/Rails. RubySage targets maintained application stacks. Older gem releases remain available for legacy Rails applications.

Compatibility

  • Ruby: 3.2+
  • Rails: 7.1+
  • Database: anything ActiveRecord supports (PostgreSQL, MySQL, SQLite tested).

CI tests across the matrix. See .github/workflows/ci.yml.

Roadmap

  • v1.5: streaming responses (SSE), Propshaft asset pipeline, optional pgvector embeddings.
  • v2: hosted RubySage Cloud — shared snapshots across dev/prod environments, no token spend on your end.

Contributing

Pull requests welcome. Run bundle exec rspec and bundle exec rubocop before opening one.

License

MIT. See LICENSE.


Built by Lanier, an applied AI studio. More free tools at lanierdev.com/tools.

About

Rails engine for agent-ready codebase retrieval, grounded answers, tool loops, and safe database queries

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

RubySage

Gem VersionCILicense: MIT

Chat with your Rails codebase. Drop in a helper, scan your app, ask questions, get answers grounded in your actual code with citations.

Why

You've been on a Rails team for six months and nobody can explain how the subscription billing flow works because Steve left in March. You ask the codebase. RubySage answers — with file citations.

How it works

  1. Scan: a daily rake task walks your codebase and produces per-file artifacts — small structured records with summaries, public symbols, and route mappings. Stored in your app's database. Secret values redacted.
  2. Retrieve: when someone asks a question, RubySage retrieves the most relevant artifacts (keyword + symbol matching, page-context boosted).
  3. Answer: the relevant artifacts go into a prompt with the question. The LLM answers, citing the specific files and classes it used.

No giant snapshot stuffed into every prompt. No invented file paths. Just retrieval grounded in your actual code.

Install

Add to your Gemfile:

gem"ruby_sage"

Then:

bundle install
rails generate ruby_sage:install
rails db:migrate

Edit config/initializers/ruby_sage.rb to set your provider and API key.

Quickstart

# config/initializers/ruby_sage.rbRubySage.configuredo |config|
config.provider=:anthropicconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.auth_check=->(c){c.current_user&.admin?}end

Run your first scan:

bundle exec rake ruby_sage:scan

Drop the widget into your layout:

<%# app/views/layouts/application.html.erb %><body><%=yield%><%=ruby_sage_widget%></body>

Open your app, click the floating button, ask "what does the PostsController do?". Done.

Providers

Anthropic (recommended)

Uses prompt caching from day 1. The retrieved-artifact context block carries a cache_control: ephemeral marker — subsequent questions within ~5 minutes hit the cache and cost roughly 10× less.

config.provider=:anthropicconfig.model="claude-sonnet-4-6"config.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)

OpenAI

Works fine. No prompt caching in V1.

config.provider=:openaiconfig.model="gpt-4.1"config.api_key=ENV.fetch("OPENAI_API_KEY",nil)

Running scans

In development

Run manually after material code changes:

bundle exec rake ruby_sage:scan

Files unchanged since the previous scan are skipped via digest cache, so re-scans are cheap.

Use your local coding agent (no API key needed)

If you already use Claude Code, Codex, or Cursor, your agent can produce the summaries instead of RubySage paying for them. Two-step flow:

bundle exec rake ruby_sage:scan:plan

Writes tmp/ruby_sage/manifest.json (one entry per scanned file, with redacted contents inlined) and tmp/ruby_sage/INSTRUCTIONS.md. Tell your agent:

Read tmp/ruby_sage/INSTRUCTIONS.md and follow it.

The agent writes tmp/ruby_sage/summaries.json. Then:

bundle exec rake ruby_sage:scan:apply

A new completed Scan lands in your DB. Files unchanged since a prior scan reuse their cached summary, so the agent only summarizes what actually changed. Your ANTHROPIC_API_KEY / OPENAI_API_KEY is never used here — the cost lives in your existing agent subscription.

In production

Daily cron is the typical pattern:

# config/schedule.rb (using the `whenever` gem) — or your Heroku Scheduler / GitHub Actions cronevery:day,at: "4am"dorake"ruby_sage:scan"end

Pre-bake in CI, ship to prod

Expensive scans don't have to run in production. Bake them in CI:

# .github/workflows/snapshot.yml
- run: bundle exec rake ruby_sage:scan
- run: bundle exec rake ruby_sage:export_artifacts > artifacts.json
- uses: actions/upload-artifact@v4with:
name: ruby_sage_snapshotpath: artifacts.json

Then on deploy:

bundle exec rake ruby_sage:import_artifacts < artifacts.json

Production skips LLM summarization entirely — zero token spend on scans.

The agent-driven flow ships to production the same way. Run scan:plan + agent loop locally or in CI, then check the resulting manifest.json + summaries.json into a deploy artifact and run scan:apply on prod with MANIFEST=... and SUMMARIES=... pointing at the shipped files.

Cost

Rough numbers, USD, with Anthropic at current Sonnet pricing. Your mileage will vary by codebase size and question patterns.

OperationFrequencyCost
Initial full scan (200-file Rails app)once~$0.10–$0.50
Daily incremental scan (50 changed files)per day~$0.05–$0.20
User question (with prompt-cache hit)per question~$0.01–$0.05
User question (cache miss / first of session)per question~$0.10–$0.30

Pre-baking in CI moves the daily-scan cost off your production token budget entirely.

AI agent integration

RubySage's retrieval layer is a public Ruby API, not just a widget. Wire it into Cursor, Claude Code, Codex, or your own coding agent to give the model indexed knowledge instead of dumping the whole codebase into context.

result=RubySage.context_for("how does subscription billing work?")result[:artifacts]# => relevant Artifact recordsresult[:citations]# => [{path:, kind:, score:, snippet:}, ...]result[:scan_id]# => Scan id used

Or hit the JSON endpoint:

curl -X POST https://your-app.com/ruby_sage/internal/retrieve \
-H "Content-Type: application/json" \
-d '{"query":"subscription billing"}'

A typical Cursor / Claude Code session can spend 50–200K input tokens orienting the model to a Rails codebase before any real work happens. Swap that for a 3K-token retrieval call and your dev token bill drops by an order of magnitude.

Generate onboarding docs

bundle exec rake ruby_sage:onboard

Writes two files based on the latest scan:

  • docs/ONBOARDING.md — for human developers joining the team. Tech stack, data model, key workflows, where to start reading, gotchas.
  • docs/AGENT_PRIMER.md — for AI coding agents. App in one sentence, stack, domain model, service layer, patterns to follow / what NOT to do. Kept under 600 words so it fits any agent's context prefix.

A new developer or AI agent runs this once and has structured context in under a minute.

CLI chat

bundle exec rake "ruby_sage:ask[how does authentication work?]"# or
QUERY="how does authentication work?" bundle exec rake ruby_sage:ask

Prints the answer + source file citations to STDOUT. Useful for shell-based AI agents that need a quick lookup without opening a browser.

Admin dashboard

Visit /ruby_sage/admin/scans for scan history, artifact counts by kind, and a "Scan now" button. Visit /ruby_sage/admin/artifacts to browse the indexed file list with summaries and public symbols. Both go through the same auth gate as the chat endpoint.

Configuration reference

RubySage.configuredo |config|
# Providerconfig.provider=:anthropic# :anthropic | :openaiconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.summarization_model="claude-haiku-4-5"# Authorization (server-side, enforced before_action)config.scope=:admin# :admin | :signed_in | :public_rate_limitedconfig.auth_check=->(c){c.current_user&.admin?}# Mode (shapes prompt + filters artifacts by audience)config.mode=:developer# :developer | :admin | :user# Audience scopingconfig.user_facing_paths=["app/views/help/**/*"]# additively tag :userconfig.audience_for=->(attrs){ ... }# full override# Admin database queries (read-only SELECT via tool loop)config.enable_database_queries=falseconfig.query_scope=->(c){"organization_id = #{c.current_user.organization_id}"}config.query_connection=->(_c){ReadOnlyDatabase.connection}config.max_query_rows=100config.query_timeout_ms=5_000config.tool_loop_max_iterations=5# Optional CSP nonce hookconfig.csp_nonce=->(c){c.content_security_policy_nonce}# Scannerconfig.scanner_include=[...]# see lib/ruby_sage/configuration.rb for defaultsconfig.scanner_exclude=[...]config.scan_retention=7# HTTPconfig.request_timeout=30config.max_retries=2end

Security

  • Server-side auth. Every chat / retrieve request goes through before_action :authorize_ruby_sage!. The widget UI is just UX; the endpoint is the gate.
  • Secret redaction. YAML values for keys matching (api_key|secret|password|token|access_key|private_key|client_secret) are replaced with [REDACTED] at scan time. ENV[...] symbol references are preserved (the model needs to know which dependencies exist), but values never are. credentials.yml.enc is excluded entirely.
  • Provider data policies. Anthropic and OpenAI receive your code summaries (not raw source by default) when you ask questions. Read each provider's data retention policies before scanning sensitive code.
  • Supported Ruby/Rails. RubySage targets maintained application stacks. Older gem releases remain available for legacy Rails applications.

Compatibility

  • Ruby: 3.2+
  • Rails: 7.1+
  • Database: anything ActiveRecord supports (PostgreSQL, MySQL, SQLite tested).

CI tests across the matrix. See .github/workflows/ci.yml.

Roadmap

  • v1.5: streaming responses (SSE), Propshaft asset pipeline, optional pgvector embeddings.
  • v2: hosted RubySage Cloud — shared snapshots across dev/prod environments, no token spend on your end.

Contributing

Pull requests welcome. Run bundle exec rspec and bundle exec rubocop before opening one.

License

MIT. See LICENSE.


Built by Lanier, an applied AI studio. More free tools at lanierdev.com/tools.

About

Rails engine for agent-ready codebase retrieval, grounded answers, tool loops, and safe database queries

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RubySage

Gem VersionCILicense: MIT

Chat with your Rails codebase. Drop in a helper, scan your app, ask questions, get answers grounded in your actual code with citations.

Why

You've been on a Rails team for six months and nobody can explain how the subscription billing flow works because Steve left in March. You ask the codebase. RubySage answers — with file citations.

How it works

  1. Scan: a daily rake task walks your codebase and produces per-file artifacts — small structured records with summaries, public symbols, and route mappings. Stored in your app's database. Secret values redacted.
  2. Retrieve: when someone asks a question, RubySage retrieves the most relevant artifacts (keyword + symbol matching, page-context boosted).
  3. Answer: the relevant artifacts go into a prompt with the question. The LLM answers, citing the specific files and classes it used.

No giant snapshot stuffed into every prompt. No invented file paths. Just retrieval grounded in your actual code.

Install

Add to your Gemfile:

gem"ruby_sage"

Then:

bundle install
rails generate ruby_sage:install
rails db:migrate

Edit config/initializers/ruby_sage.rb to set your provider and API key.

Quickstart

# config/initializers/ruby_sage.rbRubySage.configuredo |config|
config.provider=:anthropicconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.auth_check=->(c){c.current_user&.admin?}end

Run your first scan:

bundle exec rake ruby_sage:scan

Drop the widget into your layout:

<%# app/views/layouts/application.html.erb %><body><%=yield%><%=ruby_sage_widget%></body>

Open your app, click the floating button, ask "what does the PostsController do?". Done.

Providers

Anthropic (recommended)

Uses prompt caching from day 1. The retrieved-artifact context block carries a cache_control: ephemeral marker — subsequent questions within ~5 minutes hit the cache and cost roughly 10× less.

config.provider=:anthropicconfig.model="claude-sonnet-4-6"config.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)

OpenAI

Works fine. No prompt caching in V1.

config.provider=:openaiconfig.model="gpt-4.1"config.api_key=ENV.fetch("OPENAI_API_KEY",nil)

Running scans

In development

Run manually after material code changes:

bundle exec rake ruby_sage:scan

Files unchanged since the previous scan are skipped via digest cache, so re-scans are cheap.

Use your local coding agent (no API key needed)

If you already use Claude Code, Codex, or Cursor, your agent can produce the summaries instead of RubySage paying for them. Two-step flow:

bundle exec rake ruby_sage:scan:plan

Writes tmp/ruby_sage/manifest.json (one entry per scanned file, with redacted contents inlined) and tmp/ruby_sage/INSTRUCTIONS.md. Tell your agent:

Read tmp/ruby_sage/INSTRUCTIONS.md and follow it.

The agent writes tmp/ruby_sage/summaries.json. Then:

bundle exec rake ruby_sage:scan:apply

A new completed Scan lands in your DB. Files unchanged since a prior scan reuse their cached summary, so the agent only summarizes what actually changed. Your ANTHROPIC_API_KEY / OPENAI_API_KEY is never used here — the cost lives in your existing agent subscription.

In production

Daily cron is the typical pattern:

# config/schedule.rb (using the `whenever` gem) — or your Heroku Scheduler / GitHub Actions cronevery:day,at: "4am"dorake"ruby_sage:scan"end

Pre-bake in CI, ship to prod

Expensive scans don't have to run in production. Bake them in CI:

# .github/workflows/snapshot.yml
- run: bundle exec rake ruby_sage:scan
- run: bundle exec rake ruby_sage:export_artifacts > artifacts.json
- uses: actions/upload-artifact@v4with:
name: ruby_sage_snapshotpath: artifacts.json

Then on deploy:

bundle exec rake ruby_sage:import_artifacts < artifacts.json

Production skips LLM summarization entirely — zero token spend on scans.

The agent-driven flow ships to production the same way. Run scan:plan + agent loop locally or in CI, then check the resulting manifest.json + summaries.json into a deploy artifact and run scan:apply on prod with MANIFEST=... and SUMMARIES=... pointing at the shipped files.

Cost

Rough numbers, USD, with Anthropic at current Sonnet pricing. Your mileage will vary by codebase size and question patterns.

OperationFrequencyCost
Initial full scan (200-file Rails app)once~$0.10–$0.50
Daily incremental scan (50 changed files)per day~$0.05–$0.20
User question (with prompt-cache hit)per question~$0.01–$0.05
User question (cache miss / first of session)per question~$0.10–$0.30

Pre-baking in CI moves the daily-scan cost off your production token budget entirely.

AI agent integration

RubySage's retrieval layer is a public Ruby API, not just a widget. Wire it into Cursor, Claude Code, Codex, or your own coding agent to give the model indexed knowledge instead of dumping the whole codebase into context.

result=RubySage.context_for("how does subscription billing work?")result[:artifacts]# => relevant Artifact recordsresult[:citations]# => [{path:, kind:, score:, snippet:}, ...]result[:scan_id]# => Scan id used

Or hit the JSON endpoint:

curl -X POST https://your-app.com/ruby_sage/internal/retrieve \
-H "Content-Type: application/json" \
-d '{"query":"subscription billing"}'

A typical Cursor / Claude Code session can spend 50–200K input tokens orienting the model to a Rails codebase before any real work happens. Swap that for a 3K-token retrieval call and your dev token bill drops by an order of magnitude.

Generate onboarding docs

bundle exec rake ruby_sage:onboard

Writes two files based on the latest scan:

  • docs/ONBOARDING.md — for human developers joining the team. Tech stack, data model, key workflows, where to start reading, gotchas.
  • docs/AGENT_PRIMER.md — for AI coding agents. App in one sentence, stack, domain model, service layer, patterns to follow / what NOT to do. Kept under 600 words so it fits any agent's context prefix.

A new developer or AI agent runs this once and has structured context in under a minute.

CLI chat

bundle exec rake "ruby_sage:ask[how does authentication work?]"# or
QUERY="how does authentication work?" bundle exec rake ruby_sage:ask

Prints the answer + source file citations to STDOUT. Useful for shell-based AI agents that need a quick lookup without opening a browser.

Admin dashboard

Visit /ruby_sage/admin/scans for scan history, artifact counts by kind, and a "Scan now" button. Visit /ruby_sage/admin/artifacts to browse the indexed file list with summaries and public symbols. Both go through the same auth gate as the chat endpoint.

Configuration reference

RubySage.configuredo |config|
# Providerconfig.provider=:anthropic# :anthropic | :openaiconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.summarization_model="claude-haiku-4-5"# Authorization (server-side, enforced before_action)config.scope=:admin# :admin | :signed_in | :public_rate_limitedconfig.auth_check=->(c){c.current_user&.admin?}# Mode (shapes prompt + filters artifacts by audience)config.mode=:developer# :developer | :admin | :user# Audience scopingconfig.user_facing_paths=["app/views/help/**/*"]# additively tag :userconfig.audience_for=->(attrs){ ... }# full override# Admin database queries (read-only SELECT via tool loop)config.enable_database_queries=falseconfig.query_scope=->(c){"organization_id = #{c.current_user.organization_id}"}config.query_connection=->(_c){ReadOnlyDatabase.connection}config.max_query_rows=100config.query_timeout_ms=5_000config.tool_loop_max_iterations=5# Optional CSP nonce hookconfig.csp_nonce=->(c){c.content_security_policy_nonce}# Scannerconfig.scanner_include=[...]# see lib/ruby_sage/configuration.rb for defaultsconfig.scanner_exclude=[...]config.scan_retention=7# HTTPconfig.request_timeout=30config.max_retries=2end

Security

  • Server-side auth. Every chat / retrieve request goes through before_action :authorize_ruby_sage!. The widget UI is just UX; the endpoint is the gate.
  • Secret redaction. YAML values for keys matching (api_key|secret|password|token|access_key|private_key|client_secret) are replaced with [REDACTED] at scan time. ENV[...] symbol references are preserved (the model needs to know which dependencies exist), but values never are. credentials.yml.enc is excluded entirely.
  • Provider data policies. Anthropic and OpenAI receive your code summaries (not raw source by default) when you ask questions. Read each provider's data retention policies before scanning sensitive code.
  • Supported Ruby/Rails. RubySage targets maintained application stacks. Older gem releases remain available for legacy Rails applications.

Compatibility

  • Ruby: 3.2+
  • Rails: 7.1+
  • Database: anything ActiveRecord supports (PostgreSQL, MySQL, SQLite tested).

CI tests across the matrix. See .github/workflows/ci.yml.

Roadmap

  • v1.5: streaming responses (SSE), Propshaft asset pipeline, optional pgvector embeddings.
  • v2: hosted RubySage Cloud — shared snapshots across dev/prod environments, no token spend on your end.

Contributing

Pull requests welcome. Run bundle exec rspec and bundle exec rubocop before opening one.

License

MIT. See LICENSE.


Built by Lanier, an applied AI studio. More free tools at lanierdev.com/tools.

About

Rails engine for agent-ready codebase retrieval, grounded answers, tool loops, and safe database queries

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RubySage

Gem VersionCILicense: MIT

Chat with your Rails codebase. Drop in a helper, scan your app, ask questions, get answers grounded in your actual code with citations.

Why

You've been on a Rails team for six months and nobody can explain how the subscription billing flow works because Steve left in March. You ask the codebase. RubySage answers — with file citations.

How it works

  1. Scan: a daily rake task walks your codebase and produces per-file artifacts — small structured records with summaries, public symbols, and route mappings. Stored in your app's database. Secret values redacted.
  2. Retrieve: when someone asks a question, RubySage retrieves the most relevant artifacts (keyword + symbol matching, page-context boosted).
  3. Answer: the relevant artifacts go into a prompt with the question. The LLM answers, citing the specific files and classes it used.

No giant snapshot stuffed into every prompt. No invented file paths. Just retrieval grounded in your actual code.

Install

Add to your Gemfile:

gem"ruby_sage"

Then:

bundle install
rails generate ruby_sage:install
rails db:migrate

Edit config/initializers/ruby_sage.rb to set your provider and API key.

Quickstart

# config/initializers/ruby_sage.rbRubySage.configuredo |config|
config.provider=:anthropicconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.auth_check=->(c){c.current_user&.admin?}end

Run your first scan:

bundle exec rake ruby_sage:scan

Drop the widget into your layout:

<%# app/views/layouts/application.html.erb %><body><%=yield%><%=ruby_sage_widget%></body>

Open your app, click the floating button, ask "what does the PostsController do?". Done.

Providers

Anthropic (recommended)

Uses prompt caching from day 1. The retrieved-artifact context block carries a cache_control: ephemeral marker — subsequent questions within ~5 minutes hit the cache and cost roughly 10× less.

config.provider=:anthropicconfig.model="claude-sonnet-4-6"config.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)

OpenAI

Works fine. No prompt caching in V1.

config.provider=:openaiconfig.model="gpt-4.1"config.api_key=ENV.fetch("OPENAI_API_KEY",nil)

Running scans

In development

Run manually after material code changes:

bundle exec rake ruby_sage:scan

Files unchanged since the previous scan are skipped via digest cache, so re-scans are cheap.

Use your local coding agent (no API key needed)

If you already use Claude Code, Codex, or Cursor, your agent can produce the summaries instead of RubySage paying for them. Two-step flow:

bundle exec rake ruby_sage:scan:plan

Writes tmp/ruby_sage/manifest.json (one entry per scanned file, with redacted contents inlined) and tmp/ruby_sage/INSTRUCTIONS.md. Tell your agent:

Read tmp/ruby_sage/INSTRUCTIONS.md and follow it.

The agent writes tmp/ruby_sage/summaries.json. Then:

bundle exec rake ruby_sage:scan:apply

A new completed Scan lands in your DB. Files unchanged since a prior scan reuse their cached summary, so the agent only summarizes what actually changed. Your ANTHROPIC_API_KEY / OPENAI_API_KEY is never used here — the cost lives in your existing agent subscription.

In production

Daily cron is the typical pattern:

# config/schedule.rb (using the `whenever` gem) — or your Heroku Scheduler / GitHub Actions cronevery:day,at: "4am"dorake"ruby_sage:scan"end

Pre-bake in CI, ship to prod

Expensive scans don't have to run in production. Bake them in CI:

# .github/workflows/snapshot.yml
- run: bundle exec rake ruby_sage:scan
- run: bundle exec rake ruby_sage:export_artifacts > artifacts.json
- uses: actions/upload-artifact@v4with:
name: ruby_sage_snapshotpath: artifacts.json

Then on deploy:

bundle exec rake ruby_sage:import_artifacts < artifacts.json

Production skips LLM summarization entirely — zero token spend on scans.

The agent-driven flow ships to production the same way. Run scan:plan + agent loop locally or in CI, then check the resulting manifest.json + summaries.json into a deploy artifact and run scan:apply on prod with MANIFEST=... and SUMMARIES=... pointing at the shipped files.

Cost

Rough numbers, USD, with Anthropic at current Sonnet pricing. Your mileage will vary by codebase size and question patterns.

OperationFrequencyCost
Initial full scan (200-file Rails app)once~$0.10–$0.50
Daily incremental scan (50 changed files)per day~$0.05–$0.20
User question (with prompt-cache hit)per question~$0.01–$0.05
User question (cache miss / first of session)per question~$0.10–$0.30

Pre-baking in CI moves the daily-scan cost off your production token budget entirely.

AI agent integration

RubySage's retrieval layer is a public Ruby API, not just a widget. Wire it into Cursor, Claude Code, Codex, or your own coding agent to give the model indexed knowledge instead of dumping the whole codebase into context.

result=RubySage.context_for("how does subscription billing work?")result[:artifacts]# => relevant Artifact recordsresult[:citations]# => [{path:, kind:, score:, snippet:}, ...]result[:scan_id]# => Scan id used

Or hit the JSON endpoint:

curl -X POST https://your-app.com/ruby_sage/internal/retrieve \
-H "Content-Type: application/json" \
-d '{"query":"subscription billing"}'

A typical Cursor / Claude Code session can spend 50–200K input tokens orienting the model to a Rails codebase before any real work happens. Swap that for a 3K-token retrieval call and your dev token bill drops by an order of magnitude.

Generate onboarding docs

bundle exec rake ruby_sage:onboard

Writes two files based on the latest scan:

  • docs/ONBOARDING.md — for human developers joining the team. Tech stack, data model, key workflows, where to start reading, gotchas.
  • docs/AGENT_PRIMER.md — for AI coding agents. App in one sentence, stack, domain model, service layer, patterns to follow / what NOT to do. Kept under 600 words so it fits any agent's context prefix.

A new developer or AI agent runs this once and has structured context in under a minute.

CLI chat

bundle exec rake "ruby_sage:ask[how does authentication work?]"# or
QUERY="how does authentication work?" bundle exec rake ruby_sage:ask

Prints the answer + source file citations to STDOUT. Useful for shell-based AI agents that need a quick lookup without opening a browser.

Admin dashboard

Visit /ruby_sage/admin/scans for scan history, artifact counts by kind, and a "Scan now" button. Visit /ruby_sage/admin/artifacts to browse the indexed file list with summaries and public symbols. Both go through the same auth gate as the chat endpoint.

Configuration reference

RubySage.configuredo |config|
# Providerconfig.provider=:anthropic# :anthropic | :openaiconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.summarization_model="claude-haiku-4-5"# Authorization (server-side, enforced before_action)config.scope=:admin# :admin | :signed_in | :public_rate_limitedconfig.auth_check=->(c){c.current_user&.admin?}# Mode (shapes prompt + filters artifacts by audience)config.mode=:developer# :developer | :admin | :user# Audience scopingconfig.user_facing_paths=["app/views/help/**/*"]# additively tag :userconfig.audience_for=->(attrs){ ... }# full override# Admin database queries (read-only SELECT via tool loop)config.enable_database_queries=falseconfig.query_scope=->(c){"organization_id = #{c.current_user.organization_id}"}config.query_connection=->(_c){ReadOnlyDatabase.connection}config.max_query_rows=100config.query_timeout_ms=5_000config.tool_loop_max_iterations=5# Optional CSP nonce hookconfig.csp_nonce=->(c){c.content_security_policy_nonce}# Scannerconfig.scanner_include=[...]# see lib/ruby_sage/configuration.rb for defaultsconfig.scanner_exclude=[...]config.scan_retention=7# HTTPconfig.request_timeout=30config.max_retries=2end

Security

  • Server-side auth. Every chat / retrieve request goes through before_action :authorize_ruby_sage!. The widget UI is just UX; the endpoint is the gate.
  • Secret redaction. YAML values for keys matching (api_key|secret|password|token|access_key|private_key|client_secret) are replaced with [REDACTED] at scan time. ENV[...] symbol references are preserved (the model needs to know which dependencies exist), but values never are. credentials.yml.enc is excluded entirely.
  • Provider data policies. Anthropic and OpenAI receive your code summaries (not raw source by default) when you ask questions. Read each provider's data retention policies before scanning sensitive code.
  • Supported Ruby/Rails. RubySage targets maintained application stacks. Older gem releases remain available for legacy Rails applications.

Compatibility

  • Ruby: 3.2+
  • Rails: 7.1+
  • Database: anything ActiveRecord supports (PostgreSQL, MySQL, SQLite tested).

CI tests across the matrix. See .github/workflows/ci.yml.

Roadmap

  • v1.5: streaming responses (SSE), Propshaft asset pipeline, optional pgvector embeddings.
  • v2: hosted RubySage Cloud — shared snapshots across dev/prod environments, no token spend on your end.

Contributing

Pull requests welcome. Run bundle exec rspec and bundle exec rubocop before opening one.

License

MIT. See LICENSE.


Built by Lanier, an applied AI studio. More free tools at lanierdev.com/tools.

About

Rails engine for agent-ready codebase retrieval, grounded answers, tool loops, and safe database queries

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

RubySage

Gem VersionCILicense: MIT

Chat with your Rails codebase. Drop in a helper, scan your app, ask questions, get answers grounded in your actual code with citations.

Why

You've been on a Rails team for six months and nobody can explain how the subscription billing flow works because Steve left in March. You ask the codebase. RubySage answers — with file citations.

How it works

  1. Scan: a daily rake task walks your codebase and produces per-file artifacts — small structured records with summaries, public symbols, and route mappings. Stored in your app's database. Secret values redacted.
  2. Retrieve: when someone asks a question, RubySage retrieves the most relevant artifacts (keyword + symbol matching, page-context boosted).
  3. Answer: the relevant artifacts go into a prompt with the question. The LLM answers, citing the specific files and classes it used.

No giant snapshot stuffed into every prompt. No invented file paths. Just retrieval grounded in your actual code.

Install

Add to your Gemfile:

gem"ruby_sage"

Then:

bundle install
rails generate ruby_sage:install
rails db:migrate

Edit config/initializers/ruby_sage.rb to set your provider and API key.

Quickstart

# config/initializers/ruby_sage.rbRubySage.configuredo |config|
config.provider=:anthropicconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.auth_check=->(c){c.current_user&.admin?}end

Run your first scan:

bundle exec rake ruby_sage:scan

Drop the widget into your layout:

<%# app/views/layouts/application.html.erb %><body><%=yield%><%=ruby_sage_widget%></body>

Open your app, click the floating button, ask "what does the PostsController do?". Done.

Providers

Anthropic (recommended)

Uses prompt caching from day 1. The retrieved-artifact context block carries a cache_control: ephemeral marker — subsequent questions within ~5 minutes hit the cache and cost roughly 10× less.

config.provider=:anthropicconfig.model="claude-sonnet-4-6"config.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)

OpenAI

Works fine. No prompt caching in V1.

config.provider=:openaiconfig.model="gpt-4.1"config.api_key=ENV.fetch("OPENAI_API_KEY",nil)

Running scans

In development

Run manually after material code changes:

bundle exec rake ruby_sage:scan

Files unchanged since the previous scan are skipped via digest cache, so re-scans are cheap.

Use your local coding agent (no API key needed)

If you already use Claude Code, Codex, or Cursor, your agent can produce the summaries instead of RubySage paying for them. Two-step flow:

bundle exec rake ruby_sage:scan:plan

Writes tmp/ruby_sage/manifest.json (one entry per scanned file, with redacted contents inlined) and tmp/ruby_sage/INSTRUCTIONS.md. Tell your agent:

Read tmp/ruby_sage/INSTRUCTIONS.md and follow it.

The agent writes tmp/ruby_sage/summaries.json. Then:

bundle exec rake ruby_sage:scan:apply

A new completed Scan lands in your DB. Files unchanged since a prior scan reuse their cached summary, so the agent only summarizes what actually changed. Your ANTHROPIC_API_KEY / OPENAI_API_KEY is never used here — the cost lives in your existing agent subscription.

In production

Daily cron is the typical pattern:

# config/schedule.rb (using the `whenever` gem) — or your Heroku Scheduler / GitHub Actions cronevery:day,at: "4am"dorake"ruby_sage:scan"end

Pre-bake in CI, ship to prod

Expensive scans don't have to run in production. Bake them in CI:

# .github/workflows/snapshot.yml
- run: bundle exec rake ruby_sage:scan
- run: bundle exec rake ruby_sage:export_artifacts > artifacts.json
- uses: actions/upload-artifact@v4with:
name: ruby_sage_snapshotpath: artifacts.json

Then on deploy:

bundle exec rake ruby_sage:import_artifacts < artifacts.json

Production skips LLM summarization entirely — zero token spend on scans.

The agent-driven flow ships to production the same way. Run scan:plan + agent loop locally or in CI, then check the resulting manifest.json + summaries.json into a deploy artifact and run scan:apply on prod with MANIFEST=... and SUMMARIES=... pointing at the shipped files.

Cost

Rough numbers, USD, with Anthropic at current Sonnet pricing. Your mileage will vary by codebase size and question patterns.

OperationFrequencyCost
Initial full scan (200-file Rails app)once~$0.10–$0.50
Daily incremental scan (50 changed files)per day~$0.05–$0.20
User question (with prompt-cache hit)per question~$0.01–$0.05
User question (cache miss / first of session)per question~$0.10–$0.30

Pre-baking in CI moves the daily-scan cost off your production token budget entirely.

AI agent integration

RubySage's retrieval layer is a public Ruby API, not just a widget. Wire it into Cursor, Claude Code, Codex, or your own coding agent to give the model indexed knowledge instead of dumping the whole codebase into context.

result=RubySage.context_for("how does subscription billing work?")result[:artifacts]# => relevant Artifact recordsresult[:citations]# => [{path:, kind:, score:, snippet:}, ...]result[:scan_id]# => Scan id used

Or hit the JSON endpoint:

curl -X POST https://your-app.com/ruby_sage/internal/retrieve \
-H "Content-Type: application/json" \
-d '{"query":"subscription billing"}'

A typical Cursor / Claude Code session can spend 50–200K input tokens orienting the model to a Rails codebase before any real work happens. Swap that for a 3K-token retrieval call and your dev token bill drops by an order of magnitude.

Generate onboarding docs

bundle exec rake ruby_sage:onboard

Writes two files based on the latest scan:

  • docs/ONBOARDING.md — for human developers joining the team. Tech stack, data model, key workflows, where to start reading, gotchas.
  • docs/AGENT_PRIMER.md — for AI coding agents. App in one sentence, stack, domain model, service layer, patterns to follow / what NOT to do. Kept under 600 words so it fits any agent's context prefix.

A new developer or AI agent runs this once and has structured context in under a minute.

CLI chat

bundle exec rake "ruby_sage:ask[how does authentication work?]"# or
QUERY="how does authentication work?" bundle exec rake ruby_sage:ask

Prints the answer + source file citations to STDOUT. Useful for shell-based AI agents that need a quick lookup without opening a browser.

Admin dashboard

Visit /ruby_sage/admin/scans for scan history, artifact counts by kind, and a "Scan now" button. Visit /ruby_sage/admin/artifacts to browse the indexed file list with summaries and public symbols. Both go through the same auth gate as the chat endpoint.

Configuration reference

RubySage.configuredo |config|
# Providerconfig.provider=:anthropic# :anthropic | :openaiconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.summarization_model="claude-haiku-4-5"# Authorization (server-side, enforced before_action)config.scope=:admin# :admin | :signed_in | :public_rate_limitedconfig.auth_check=->(c){c.current_user&.admin?}# Mode (shapes prompt + filters artifacts by audience)config.mode=:developer# :developer | :admin | :user# Audience scopingconfig.user_facing_paths=["app/views/help/**/*"]# additively tag :userconfig.audience_for=->(attrs){ ... }# full override# Admin database queries (read-only SELECT via tool loop)config.enable_database_queries=falseconfig.query_scope=->(c){"organization_id = #{c.current_user.organization_id}"}config.query_connection=->(_c){ReadOnlyDatabase.connection}config.max_query_rows=100config.query_timeout_ms=5_000config.tool_loop_max_iterations=5# Optional CSP nonce hookconfig.csp_nonce=->(c){c.content_security_policy_nonce}# Scannerconfig.scanner_include=[...]# see lib/ruby_sage/configuration.rb for defaultsconfig.scanner_exclude=[...]config.scan_retention=7# HTTPconfig.request_timeout=30config.max_retries=2end

Security

  • Server-side auth. Every chat / retrieve request goes through before_action :authorize_ruby_sage!. The widget UI is just UX; the endpoint is the gate.
  • Secret redaction. YAML values for keys matching (api_key|secret|password|token|access_key|private_key|client_secret) are replaced with [REDACTED] at scan time. ENV[...] symbol references are preserved (the model needs to know which dependencies exist), but values never are. credentials.yml.enc is excluded entirely.
  • Provider data policies. Anthropic and OpenAI receive your code summaries (not raw source by default) when you ask questions. Read each provider's data retention policies before scanning sensitive code.
  • Supported Ruby/Rails. RubySage targets maintained application stacks. Older gem releases remain available for legacy Rails applications.

Compatibility

  • Ruby: 3.2+
  • Rails: 7.1+
  • Database: anything ActiveRecord supports (PostgreSQL, MySQL, SQLite tested).

CI tests across the matrix. See .github/workflows/ci.yml.

Roadmap

  • v1.5: streaming responses (SSE), Propshaft asset pipeline, optional pgvector embeddings.
  • v2: hosted RubySage Cloud — shared snapshots across dev/prod environments, no token spend on your end.

Contributing

Pull requests welcome. Run bundle exec rspec and bundle exec rubocop before opening one.

License

MIT. See LICENSE.


Built by Lanier, an applied AI studio. More free tools at lanierdev.com/tools.

About

Rails engine for agent-ready codebase retrieval, grounded answers, tool loops, and safe database queries

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RubySage

Gem VersionCILicense: MIT

Chat with your Rails codebase. Drop in a helper, scan your app, ask questions, get answers grounded in your actual code with citations.

Why

You've been on a Rails team for six months and nobody can explain how the subscription billing flow works because Steve left in March. You ask the codebase. RubySage answers — with file citations.

How it works

  1. Scan: a daily rake task walks your codebase and produces per-file artifacts — small structured records with summaries, public symbols, and route mappings. Stored in your app's database. Secret values redacted.
  2. Retrieve: when someone asks a question, RubySage retrieves the most relevant artifacts (keyword + symbol matching, page-context boosted).
  3. Answer: the relevant artifacts go into a prompt with the question. The LLM answers, citing the specific files and classes it used.

No giant snapshot stuffed into every prompt. No invented file paths. Just retrieval grounded in your actual code.

Install

Add to your Gemfile:

gem"ruby_sage"

Then:

bundle install
rails generate ruby_sage:install
rails db:migrate

Edit config/initializers/ruby_sage.rb to set your provider and API key.

Quickstart

# config/initializers/ruby_sage.rbRubySage.configuredo |config|
config.provider=:anthropicconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.auth_check=->(c){c.current_user&.admin?}end

Run your first scan:

bundle exec rake ruby_sage:scan

Drop the widget into your layout:

<%# app/views/layouts/application.html.erb %><body><%=yield%><%=ruby_sage_widget%></body>

Open your app, click the floating button, ask "what does the PostsController do?". Done.

Providers

Anthropic (recommended)

Uses prompt caching from day 1. The retrieved-artifact context block carries a cache_control: ephemeral marker — subsequent questions within ~5 minutes hit the cache and cost roughly 10× less.

config.provider=:anthropicconfig.model="claude-sonnet-4-6"config.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)

OpenAI

Works fine. No prompt caching in V1.

config.provider=:openaiconfig.model="gpt-4.1"config.api_key=ENV.fetch("OPENAI_API_KEY",nil)

Running scans

In development

Run manually after material code changes:

bundle exec rake ruby_sage:scan

Files unchanged since the previous scan are skipped via digest cache, so re-scans are cheap.

Use your local coding agent (no API key needed)

If you already use Claude Code, Codex, or Cursor, your agent can produce the summaries instead of RubySage paying for them. Two-step flow:

bundle exec rake ruby_sage:scan:plan

Writes tmp/ruby_sage/manifest.json (one entry per scanned file, with redacted contents inlined) and tmp/ruby_sage/INSTRUCTIONS.md. Tell your agent:

Read tmp/ruby_sage/INSTRUCTIONS.md and follow it.

The agent writes tmp/ruby_sage/summaries.json. Then:

bundle exec rake ruby_sage:scan:apply

A new completed Scan lands in your DB. Files unchanged since a prior scan reuse their cached summary, so the agent only summarizes what actually changed. Your ANTHROPIC_API_KEY / OPENAI_API_KEY is never used here — the cost lives in your existing agent subscription.

In production

Daily cron is the typical pattern:

# config/schedule.rb (using the `whenever` gem) — or your Heroku Scheduler / GitHub Actions cronevery:day,at: "4am"dorake"ruby_sage:scan"end

Pre-bake in CI, ship to prod

Expensive scans don't have to run in production. Bake them in CI:

# .github/workflows/snapshot.yml
- run: bundle exec rake ruby_sage:scan
- run: bundle exec rake ruby_sage:export_artifacts > artifacts.json
- uses: actions/upload-artifact@v4with:
name: ruby_sage_snapshotpath: artifacts.json

Then on deploy:

bundle exec rake ruby_sage:import_artifacts < artifacts.json

Production skips LLM summarization entirely — zero token spend on scans.

The agent-driven flow ships to production the same way. Run scan:plan + agent loop locally or in CI, then check the resulting manifest.json + summaries.json into a deploy artifact and run scan:apply on prod with MANIFEST=... and SUMMARIES=... pointing at the shipped files.

Cost

Rough numbers, USD, with Anthropic at current Sonnet pricing. Your mileage will vary by codebase size and question patterns.

OperationFrequencyCost
Initial full scan (200-file Rails app)once~$0.10–$0.50
Daily incremental scan (50 changed files)per day~$0.05–$0.20
User question (with prompt-cache hit)per question~$0.01–$0.05
User question (cache miss / first of session)per question~$0.10–$0.30

Pre-baking in CI moves the daily-scan cost off your production token budget entirely.

AI agent integration

RubySage's retrieval layer is a public Ruby API, not just a widget. Wire it into Cursor, Claude Code, Codex, or your own coding agent to give the model indexed knowledge instead of dumping the whole codebase into context.

result=RubySage.context_for("how does subscription billing work?")result[:artifacts]# => relevant Artifact recordsresult[:citations]# => [{path:, kind:, score:, snippet:}, ...]result[:scan_id]# => Scan id used

Or hit the JSON endpoint:

curl -X POST https://your-app.com/ruby_sage/internal/retrieve \
-H "Content-Type: application/json" \
-d '{"query":"subscription billing"}'

A typical Cursor / Claude Code session can spend 50–200K input tokens orienting the model to a Rails codebase before any real work happens. Swap that for a 3K-token retrieval call and your dev token bill drops by an order of magnitude.

Generate onboarding docs

bundle exec rake ruby_sage:onboard

Writes two files based on the latest scan:

  • docs/ONBOARDING.md — for human developers joining the team. Tech stack, data model, key workflows, where to start reading, gotchas.
  • docs/AGENT_PRIMER.md — for AI coding agents. App in one sentence, stack, domain model, service layer, patterns to follow / what NOT to do. Kept under 600 words so it fits any agent's context prefix.

A new developer or AI agent runs this once and has structured context in under a minute.

CLI chat

bundle exec rake "ruby_sage:ask[how does authentication work?]"# or
QUERY="how does authentication work?" bundle exec rake ruby_sage:ask

Prints the answer + source file citations to STDOUT. Useful for shell-based AI agents that need a quick lookup without opening a browser.

Admin dashboard

Visit /ruby_sage/admin/scans for scan history, artifact counts by kind, and a "Scan now" button. Visit /ruby_sage/admin/artifacts to browse the indexed file list with summaries and public symbols. Both go through the same auth gate as the chat endpoint.

Configuration reference

RubySage.configuredo |config|
# Providerconfig.provider=:anthropic# :anthropic | :openaiconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.summarization_model="claude-haiku-4-5"# Authorization (server-side, enforced before_action)config.scope=:admin# :admin | :signed_in | :public_rate_limitedconfig.auth_check=->(c){c.current_user&.admin?}# Mode (shapes prompt + filters artifacts by audience)config.mode=:developer# :developer | :admin | :user# Audience scopingconfig.user_facing_paths=["app/views/help/**/*"]# additively tag :userconfig.audience_for=->(attrs){ ... }# full override# Admin database queries (read-only SELECT via tool loop)config.enable_database_queries=falseconfig.query_scope=->(c){"organization_id = #{c.current_user.organization_id}"}config.query_connection=->(_c){ReadOnlyDatabase.connection}config.max_query_rows=100config.query_timeout_ms=5_000config.tool_loop_max_iterations=5# Optional CSP nonce hookconfig.csp_nonce=->(c){c.content_security_policy_nonce}# Scannerconfig.scanner_include=[...]# see lib/ruby_sage/configuration.rb for defaultsconfig.scanner_exclude=[...]config.scan_retention=7# HTTPconfig.request_timeout=30config.max_retries=2end

Security

  • Server-side auth. Every chat / retrieve request goes through before_action :authorize_ruby_sage!. The widget UI is just UX; the endpoint is the gate.
  • Secret redaction. YAML values for keys matching (api_key|secret|password|token|access_key|private_key|client_secret) are replaced with [REDACTED] at scan time. ENV[...] symbol references are preserved (the model needs to know which dependencies exist), but values never are. credentials.yml.enc is excluded entirely.
  • Provider data policies. Anthropic and OpenAI receive your code summaries (not raw source by default) when you ask questions. Read each provider's data retention policies before scanning sensitive code.
  • Supported Ruby/Rails. RubySage targets maintained application stacks. Older gem releases remain available for legacy Rails applications.

Compatibility

  • Ruby: 3.2+
  • Rails: 7.1+
  • Database: anything ActiveRecord supports (PostgreSQL, MySQL, SQLite tested).

CI tests across the matrix. See .github/workflows/ci.yml.

Roadmap

  • v1.5: streaming responses (SSE), Propshaft asset pipeline, optional pgvector embeddings.
  • v2: hosted RubySage Cloud — shared snapshots across dev/prod environments, no token spend on your end.

Contributing

Pull requests welcome. Run bundle exec rspec and bundle exec rubocop before opening one.

License

MIT. See LICENSE.


Built by Lanier, an applied AI studio. More free tools at lanierdev.com/tools.

About

Rails engine for agent-ready codebase retrieval, grounded answers, tool loops, and safe database queries

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RubySage

Gem VersionCILicense: MIT

Chat with your Rails codebase. Drop in a helper, scan your app, ask questions, get answers grounded in your actual code with citations.

Why

You've been on a Rails team for six months and nobody can explain how the subscription billing flow works because Steve left in March. You ask the codebase. RubySage answers — with file citations.

How it works

  1. Scan: a daily rake task walks your codebase and produces per-file artifacts — small structured records with summaries, public symbols, and route mappings. Stored in your app's database. Secret values redacted.
  2. Retrieve: when someone asks a question, RubySage retrieves the most relevant artifacts (keyword + symbol matching, page-context boosted).
  3. Answer: the relevant artifacts go into a prompt with the question. The LLM answers, citing the specific files and classes it used.

No giant snapshot stuffed into every prompt. No invented file paths. Just retrieval grounded in your actual code.

Install

Add to your Gemfile:

gem"ruby_sage"

Then:

bundle install
rails generate ruby_sage:install
rails db:migrate

Edit config/initializers/ruby_sage.rb to set your provider and API key.

Quickstart

# config/initializers/ruby_sage.rbRubySage.configuredo |config|
config.provider=:anthropicconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.auth_check=->(c){c.current_user&.admin?}end

Run your first scan:

bundle exec rake ruby_sage:scan

Drop the widget into your layout:

<%# app/views/layouts/application.html.erb %><body><%=yield%><%=ruby_sage_widget%></body>

Open your app, click the floating button, ask "what does the PostsController do?". Done.

Providers

Anthropic (recommended)

Uses prompt caching from day 1. The retrieved-artifact context block carries a cache_control: ephemeral marker — subsequent questions within ~5 minutes hit the cache and cost roughly 10× less.

config.provider=:anthropicconfig.model="claude-sonnet-4-6"config.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)

OpenAI

Works fine. No prompt caching in V1.

config.provider=:openaiconfig.model="gpt-4.1"config.api_key=ENV.fetch("OPENAI_API_KEY",nil)

Running scans

In development

Run manually after material code changes:

bundle exec rake ruby_sage:scan

Files unchanged since the previous scan are skipped via digest cache, so re-scans are cheap.

Use your local coding agent (no API key needed)

If you already use Claude Code, Codex, or Cursor, your agent can produce the summaries instead of RubySage paying for them. Two-step flow:

bundle exec rake ruby_sage:scan:plan

Writes tmp/ruby_sage/manifest.json (one entry per scanned file, with redacted contents inlined) and tmp/ruby_sage/INSTRUCTIONS.md. Tell your agent:

Read tmp/ruby_sage/INSTRUCTIONS.md and follow it.

The agent writes tmp/ruby_sage/summaries.json. Then:

bundle exec rake ruby_sage:scan:apply

A new completed Scan lands in your DB. Files unchanged since a prior scan reuse their cached summary, so the agent only summarizes what actually changed. Your ANTHROPIC_API_KEY / OPENAI_API_KEY is never used here — the cost lives in your existing agent subscription.

In production

Daily cron is the typical pattern:

# config/schedule.rb (using the `whenever` gem) — or your Heroku Scheduler / GitHub Actions cronevery:day,at: "4am"dorake"ruby_sage:scan"end

Pre-bake in CI, ship to prod

Expensive scans don't have to run in production. Bake them in CI:

# .github/workflows/snapshot.yml
- run: bundle exec rake ruby_sage:scan
- run: bundle exec rake ruby_sage:export_artifacts > artifacts.json
- uses: actions/upload-artifact@v4with:
name: ruby_sage_snapshotpath: artifacts.json

Then on deploy:

bundle exec rake ruby_sage:import_artifacts < artifacts.json

Production skips LLM summarization entirely — zero token spend on scans.

The agent-driven flow ships to production the same way. Run scan:plan + agent loop locally or in CI, then check the resulting manifest.json + summaries.json into a deploy artifact and run scan:apply on prod with MANIFEST=... and SUMMARIES=... pointing at the shipped files.

Cost

Rough numbers, USD, with Anthropic at current Sonnet pricing. Your mileage will vary by codebase size and question patterns.

OperationFrequencyCost
Initial full scan (200-file Rails app)once~$0.10–$0.50
Daily incremental scan (50 changed files)per day~$0.05–$0.20
User question (with prompt-cache hit)per question~$0.01–$0.05
User question (cache miss / first of session)per question~$0.10–$0.30

Pre-baking in CI moves the daily-scan cost off your production token budget entirely.

AI agent integration

RubySage's retrieval layer is a public Ruby API, not just a widget. Wire it into Cursor, Claude Code, Codex, or your own coding agent to give the model indexed knowledge instead of dumping the whole codebase into context.

result=RubySage.context_for("how does subscription billing work?")result[:artifacts]# => relevant Artifact recordsresult[:citations]# => [{path:, kind:, score:, snippet:}, ...]result[:scan_id]# => Scan id used

Or hit the JSON endpoint:

curl -X POST https://your-app.com/ruby_sage/internal/retrieve \
-H "Content-Type: application/json" \
-d '{"query":"subscription billing"}'

A typical Cursor / Claude Code session can spend 50–200K input tokens orienting the model to a Rails codebase before any real work happens. Swap that for a 3K-token retrieval call and your dev token bill drops by an order of magnitude.

Generate onboarding docs

bundle exec rake ruby_sage:onboard

Writes two files based on the latest scan:

  • docs/ONBOARDING.md — for human developers joining the team. Tech stack, data model, key workflows, where to start reading, gotchas.
  • docs/AGENT_PRIMER.md — for AI coding agents. App in one sentence, stack, domain model, service layer, patterns to follow / what NOT to do. Kept under 600 words so it fits any agent's context prefix.

A new developer or AI agent runs this once and has structured context in under a minute.

CLI chat

bundle exec rake "ruby_sage:ask[how does authentication work?]"# or
QUERY="how does authentication work?" bundle exec rake ruby_sage:ask

Prints the answer + source file citations to STDOUT. Useful for shell-based AI agents that need a quick lookup without opening a browser.

Admin dashboard

Visit /ruby_sage/admin/scans for scan history, artifact counts by kind, and a "Scan now" button. Visit /ruby_sage/admin/artifacts to browse the indexed file list with summaries and public symbols. Both go through the same auth gate as the chat endpoint.

Configuration reference

RubySage.configuredo |config|
# Providerconfig.provider=:anthropic# :anthropic | :openaiconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.summarization_model="claude-haiku-4-5"# Authorization (server-side, enforced before_action)config.scope=:admin# :admin | :signed_in | :public_rate_limitedconfig.auth_check=->(c){c.current_user&.admin?}# Mode (shapes prompt + filters artifacts by audience)config.mode=:developer# :developer | :admin | :user# Audience scopingconfig.user_facing_paths=["app/views/help/**/*"]# additively tag :userconfig.audience_for=->(attrs){ ... }# full override# Admin database queries (read-only SELECT via tool loop)config.enable_database_queries=falseconfig.query_scope=->(c){"organization_id = #{c.current_user.organization_id}"}config.query_connection=->(_c){ReadOnlyDatabase.connection}config.max_query_rows=100config.query_timeout_ms=5_000config.tool_loop_max_iterations=5# Optional CSP nonce hookconfig.csp_nonce=->(c){c.content_security_policy_nonce}# Scannerconfig.scanner_include=[...]# see lib/ruby_sage/configuration.rb for defaultsconfig.scanner_exclude=[...]config.scan_retention=7# HTTPconfig.request_timeout=30config.max_retries=2end

Security

  • Server-side auth. Every chat / retrieve request goes through before_action :authorize_ruby_sage!. The widget UI is just UX; the endpoint is the gate.
  • Secret redaction. YAML values for keys matching (api_key|secret|password|token|access_key|private_key|client_secret) are replaced with [REDACTED] at scan time. ENV[...] symbol references are preserved (the model needs to know which dependencies exist), but values never are. credentials.yml.enc is excluded entirely.
  • Provider data policies. Anthropic and OpenAI receive your code summaries (not raw source by default) when you ask questions. Read each provider's data retention policies before scanning sensitive code.
  • Supported Ruby/Rails. RubySage targets maintained application stacks. Older gem releases remain available for legacy Rails applications.

Compatibility

  • Ruby: 3.2+
  • Rails: 7.1+
  • Database: anything ActiveRecord supports (PostgreSQL, MySQL, SQLite tested).

CI tests across the matrix. See .github/workflows/ci.yml.

Roadmap

  • v1.5: streaming responses (SSE), Propshaft asset pipeline, optional pgvector embeddings.
  • v2: hosted RubySage Cloud — shared snapshots across dev/prod environments, no token spend on your end.

Contributing

Pull requests welcome. Run bundle exec rspec and bundle exec rubocop before opening one.

License

MIT. See LICENSE.


Built by Lanier, an applied AI studio. More free tools at lanierdev.com/tools.

About

Rails engine for agent-ready codebase retrieval, grounded answers, tool loops, and safe database queries

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

RubySage

Gem VersionCILicense: MIT

Chat with your Rails codebase. Drop in a helper, scan your app, ask questions, get answers grounded in your actual code with citations.

Why

You've been on a Rails team for six months and nobody can explain how the subscription billing flow works because Steve left in March. You ask the codebase. RubySage answers — with file citations.

How it works

  1. Scan: a daily rake task walks your codebase and produces per-file artifacts — small structured records with summaries, public symbols, and route mappings. Stored in your app's database. Secret values redacted.
  2. Retrieve: when someone asks a question, RubySage retrieves the most relevant artifacts (keyword + symbol matching, page-context boosted).
  3. Answer: the relevant artifacts go into a prompt with the question. The LLM answers, citing the specific files and classes it used.

No giant snapshot stuffed into every prompt. No invented file paths. Just retrieval grounded in your actual code.

Install

Add to your Gemfile:

gem"ruby_sage"

Then:

bundle install
rails generate ruby_sage:install
rails db:migrate

Edit config/initializers/ruby_sage.rb to set your provider and API key.

Quickstart

# config/initializers/ruby_sage.rbRubySage.configuredo |config|
config.provider=:anthropicconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.auth_check=->(c){c.current_user&.admin?}end

Run your first scan:

bundle exec rake ruby_sage:scan

Drop the widget into your layout:

<%# app/views/layouts/application.html.erb %><body><%=yield%><%=ruby_sage_widget%></body>

Open your app, click the floating button, ask "what does the PostsController do?". Done.

Providers

Anthropic (recommended)

Uses prompt caching from day 1. The retrieved-artifact context block carries a cache_control: ephemeral marker — subsequent questions within ~5 minutes hit the cache and cost roughly 10× less.

config.provider=:anthropicconfig.model="claude-sonnet-4-6"config.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)

OpenAI

Works fine. No prompt caching in V1.

config.provider=:openaiconfig.model="gpt-4.1"config.api_key=ENV.fetch("OPENAI_API_KEY",nil)

Running scans

In development

Run manually after material code changes:

bundle exec rake ruby_sage:scan

Files unchanged since the previous scan are skipped via digest cache, so re-scans are cheap.

Use your local coding agent (no API key needed)

If you already use Claude Code, Codex, or Cursor, your agent can produce the summaries instead of RubySage paying for them. Two-step flow:

bundle exec rake ruby_sage:scan:plan

Writes tmp/ruby_sage/manifest.json (one entry per scanned file, with redacted contents inlined) and tmp/ruby_sage/INSTRUCTIONS.md. Tell your agent:

Read tmp/ruby_sage/INSTRUCTIONS.md and follow it.

The agent writes tmp/ruby_sage/summaries.json. Then:

bundle exec rake ruby_sage:scan:apply

A new completed Scan lands in your DB. Files unchanged since a prior scan reuse their cached summary, so the agent only summarizes what actually changed. Your ANTHROPIC_API_KEY / OPENAI_API_KEY is never used here — the cost lives in your existing agent subscription.

In production

Daily cron is the typical pattern:

# config/schedule.rb (using the `whenever` gem) — or your Heroku Scheduler / GitHub Actions cronevery:day,at: "4am"dorake"ruby_sage:scan"end

Pre-bake in CI, ship to prod

Expensive scans don't have to run in production. Bake them in CI:

# .github/workflows/snapshot.yml
- run: bundle exec rake ruby_sage:scan
- run: bundle exec rake ruby_sage:export_artifacts > artifacts.json
- uses: actions/upload-artifact@v4with:
name: ruby_sage_snapshotpath: artifacts.json

Then on deploy:

bundle exec rake ruby_sage:import_artifacts < artifacts.json

Production skips LLM summarization entirely — zero token spend on scans.

The agent-driven flow ships to production the same way. Run scan:plan + agent loop locally or in CI, then check the resulting manifest.json + summaries.json into a deploy artifact and run scan:apply on prod with MANIFEST=... and SUMMARIES=... pointing at the shipped files.

Cost

Rough numbers, USD, with Anthropic at current Sonnet pricing. Your mileage will vary by codebase size and question patterns.

OperationFrequencyCost
Initial full scan (200-file Rails app)once~$0.10–$0.50
Daily incremental scan (50 changed files)per day~$0.05–$0.20
User question (with prompt-cache hit)per question~$0.01–$0.05
User question (cache miss / first of session)per question~$0.10–$0.30

Pre-baking in CI moves the daily-scan cost off your production token budget entirely.

AI agent integration

RubySage's retrieval layer is a public Ruby API, not just a widget. Wire it into Cursor, Claude Code, Codex, or your own coding agent to give the model indexed knowledge instead of dumping the whole codebase into context.

result=RubySage.context_for("how does subscription billing work?")result[:artifacts]# => relevant Artifact recordsresult[:citations]# => [{path:, kind:, score:, snippet:}, ...]result[:scan_id]# => Scan id used

Or hit the JSON endpoint:

curl -X POST https://your-app.com/ruby_sage/internal/retrieve \
-H "Content-Type: application/json" \
-d '{"query":"subscription billing"}'

A typical Cursor / Claude Code session can spend 50–200K input tokens orienting the model to a Rails codebase before any real work happens. Swap that for a 3K-token retrieval call and your dev token bill drops by an order of magnitude.

Generate onboarding docs

bundle exec rake ruby_sage:onboard

Writes two files based on the latest scan:

  • docs/ONBOARDING.md — for human developers joining the team. Tech stack, data model, key workflows, where to start reading, gotchas.
  • docs/AGENT_PRIMER.md — for AI coding agents. App in one sentence, stack, domain model, service layer, patterns to follow / what NOT to do. Kept under 600 words so it fits any agent's context prefix.

A new developer or AI agent runs this once and has structured context in under a minute.

CLI chat

bundle exec rake "ruby_sage:ask[how does authentication work?]"# or
QUERY="how does authentication work?" bundle exec rake ruby_sage:ask

Prints the answer + source file citations to STDOUT. Useful for shell-based AI agents that need a quick lookup without opening a browser.

Admin dashboard

Visit /ruby_sage/admin/scans for scan history, artifact counts by kind, and a "Scan now" button. Visit /ruby_sage/admin/artifacts to browse the indexed file list with summaries and public symbols. Both go through the same auth gate as the chat endpoint.

Configuration reference

RubySage.configuredo |config|
# Providerconfig.provider=:anthropic# :anthropic | :openaiconfig.api_key=ENV.fetch("ANTHROPIC_API_KEY",nil)config.model="claude-sonnet-4-6"config.summarization_model="claude-haiku-4-5"# Authorization (server-side, enforced before_action)config.scope=:admin# :admin | :signed_in | :public_rate_limitedconfig.auth_check=->(c){c.current_user&.admin?}# Mode (shapes prompt + filters artifacts by audience)config.mode=:developer# :developer | :admin | :user# Audience scopingconfig.user_facing_paths=["app/views/help/**/*"]# additively tag :userconfig.audience_for=->(attrs){ ... }# full override# Admin database queries (read-only SELECT via tool loop)config.enable_database_queries=falseconfig.query_scope=->(c){"organization_id = #{c.current_user.organization_id}"}config.query_connection=->(_c){ReadOnlyDatabase.connection}config.max_query_rows=100config.query_timeout_ms=5_000config.tool_loop_max_iterations=5# Optional CSP nonce hookconfig.csp_nonce=->(c){c.content_security_policy_nonce}# Scannerconfig.scanner_include=[...]# see lib/ruby_sage/configuration.rb for defaultsconfig.scanner_exclude=[...]config.scan_retention=7# HTTPconfig.request_timeout=30config.max_retries=2end

Security

  • Server-side auth. Every chat / retrieve request goes through before_action :authorize_ruby_sage!. The widget UI is just UX; the endpoint is the gate.
  • Secret redaction. YAML values for keys matching (api_key|secret|password|token|access_key|private_key|client_secret) are replaced with [REDACTED] at scan time. ENV[...] symbol references are preserved (the model needs to know which dependencies exist), but values never are. credentials.yml.enc is excluded entirely.
  • Provider data policies. Anthropic and OpenAI receive your code summaries (not raw source by default) when you ask questions. Read each provider's data retention policies before scanning sensitive code.
  • Supported Ruby/Rails. RubySage targets maintained application stacks. Older gem releases remain available for legacy Rails applications.

Compatibility

  • Ruby: 3.2+
  • Rails: 7.1+
  • Database: anything ActiveRecord supports (PostgreSQL, MySQL, SQLite tested).

CI tests across the matrix. See .github/workflows/ci.yml.

Roadmap

  • v1.5: streaming responses (SSE), Propshaft asset pipeline, optional pgvector embeddings.
  • v2: hosted RubySage Cloud — shared snapshots across dev/prod environments, no token spend on your end.

Contributing

Pull requests welcome. Run bundle exec rspec and bundle exec rubocop before opening one.

License

MIT. See LICENSE.


Built by Lanier, an applied AI studio. More free tools at lanierdev.com/tools.

About

Rails engine for agent-ready codebase retrieval, grounded answers, tool loops, and safe database queries

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages