Official Ruby client for Vision API — send an image or a PDF, describe the fields you want in plain language, get structured JSON back with a confidence level on every value.
- Website — https://visionapi.io
- Documentation — https://docs.visionapi.io
- API keys — https://app.visionapi.io/dashboard/keys
- Preset catalog — https://visionapi.io/presets
- Playground — https://visionapi.io/playground
- Support — https://support.visionapi.io · https://visionapi.io/contact-us
# Gemfilegem"vision_api"gem install vision_apiRuby 3.0+. No runtime dependencies — net/http, json and openssl are all standard
library.
require"vision_api"vision=VisionAPI.new# reads ENV["VISION_API_KEY"]res=vision.analyze(file: "invoice.pdf",preset: "invoice")res["result"]["invoice_id"]["value"]# => "A-10422"res["result"]["total"]["value"]# => 1284.5, or nil if the invoice has no totalres["credits_used"]# => 3Requests are metered in credits, per image and per selected PDF page — see pricing for current rates. Failures cost nothing: the reservation is released in full on any non-2xx, so there is no compensating logic to write.
Server-side only. There is no publishable key and no test mode — an API key is a live spending credential. Keep it in the environment (or Rails credentials), never in a repository and never in anything a browser downloads.
Responses are plain hashes with the wire's keys, so everything you already know about hashes applies. Two rules explain almost every surprise:
1. Every scalar is wrapped.{"value" => …, "confidence" => "low"|"mid"|"high"}. Read
res["result"]["total"]["value"], not res["result"]["total"].
2. A preset response contains every field of that preset — including the ones the
document does not carry, which come back as {"value" => nil, "confidence" => "low"}. A
key being present does not mean a value was found.
Line-item arrays are the one shape worth looking at twice. The array itself is not wrapped; each cell inside each row is:
{"invoice_id"=>{"value"=>"A-10422","confidence"=>"high"},"carrier"=>{"value"=>nil,"confidence"=>"low"},"line_item"=>[{"description"=>{"value"=>"Widget","confidence"=>"high"},"quantity"=>{"value"=>2,"confidence"=>"high"},"amount"=>{"value"=>25.0,"confidence"=>"mid"}}]}VisionAPI::Result covers the common readings, so you rarely have to spell that out:
includeVisionAPI::Result# or call them as VisionAPI::Result.unwrap(…)unwrap(res["result"])# => {"invoice_id" => "A-10422", "carrier" => nil, "line_item" => [{"description" => "Widget", …}]}unwrap(res["result"],drop_null: true)# only what was actually foundvalue(res["result"],"total",0)# 1284.5, or 0 when absentrows(res["result"],"line_item")# [] when the invoice has no linespresent(res["result"])# ["invoice_id", "total", "line_item"]missing(res["result"])# ["carrier", …]below_confidence(res["result"],"high")# fields to route to a humanExactly one file source per call:
vision.analyze(file: "invoice.pdf",preset: "invoice")# a pathvision.analyze(file: Pathname("invoice.pdf"),preset: "invoice")# a Pathnamevision.analyze(file: File.open("invoice.pdf","rb"),…)# an open binary IOvision.analyze(file: ["scan.png",bytes],…)# bytes + a namevision.analyze(file_url: "https://example.com/invoice.pdf",…)# a public URLvision.analyze(file_base64: encoded,…)# "data:" prefix optionalJPEG, PNG, WebP, TIFF and PDF, up to 20 MB and 50 pages. The type is detected from magic bytes — the filename is ignored.
| Keyword | Default | What it does |
|---|---|---|
preset: | — | A catalog name, or "auto" to let the API classify the file first (free). |
schema: | — | Custom fields, alone or on top of a preset. |
schema_name: | — | A schema saved in your dashboard. Excludes preset: and schema:. |
pages: | all | PDF page selection, e.g. "1-3,7". You pay for selected pages only. |
language_hint: | auto | ISO 639-1 code, e.g. "es". |
detail: | "standard" | "high" renders pages at higher resolution. Same cost, slower. |
output: | "json" | "text" returns raw OCR text instead of fields. |
include_raw_text: | false | Adds full_text, the whole transcription, alongside result. |
min_confidence: | "low" | Fields below the level come back nil, with confidence preserved. |
A schema is a flat hash: each key is a field name, each value describes what to extract. It is compiled before any credit moves, so a bad schema costs nothing.
res=vision.analyze(file: "invoice.pdf",preset: "invoice",schema: {# Plain form — the string is the description, type defaults to string."machine_serial"=>'Serial number of the machine being invoiced, without the "SN:" prefix',# Typed form."total_net"=>{"type"=>"number","description"=>"Total before tax"},"signed_on"=>{"type"=>"date","description"=>"Date the contract was signed"},# Reserved key: injects fields into every row of the preset's line-item array."line_item"=>{"lot_number"=>"The lot number printed on the line, if present"}})Field names must match ^[a-z][a-z0-9_]{0,63}$. Types are string (default), number,
boolean, date, array and object. A custom name that collides with a preset field is
a 422 schema_field_conflict — rename it, or use the preset's own field.
Descriptions are the prompt. "The invoice number exactly as printed, without the #"
extracts better than "invoice number". Say what to do when the value is missing or
ambiguous if it matters.
Reuse a combination by saving it:
vision.create_schema("our-invoices",preset: "invoice",schema: {"machine_serial"=>"…"})vision.analyze(file: "invoice.pdf",schema_name: "our-invoices")28 presets ship with the API. Fetch the catalog rather than hardcoding field names from memory — presets are versioned, and the catalog is the source of truth:
vision.presets.each{ |p| puts"#{p['name']} (#{p['kind']}) — #{p['field_count']} fields"}vision.preset("invoice")["fields"].map{ |f| f["name"]}Three ways to choose:
# 1. You know what it is.vision.analyze(file: "receipt.jpg",preset: "receipt")# 2. You don't, and you want the data anyway. Classification is free.res=vision.analyze(file: "unknown.pdf",preset: "auto")res["detection"]["preset"]# what ranres["detection"]["fallback"]# true = "shape unknown", not a matchres["detection"]["alternatives"]# the rest of the ranking, best first# 3. The *type* is the decision — routing a mixed inbox, or refusing to spend# on a 40-page PDF until you know what it is. Far cheaper than extracting.guess=vision.detect(file: "unknown.pdf")vision.analyze(file: "unknown.pdf",preset: guess["recommended"])unlessguess["fallback"]detect reads page 1 only, so an image and a 300-page PDF cost the same, and it is metered
in batches rather than per call: most calls report credits_used 0 and an occasional one
carries the charge. See pricing for the rate.
Up to 5 questions about one file, priced exactly like an extraction. The questions themselves are free.
res=vision.ask(file: "photo.jpg",questions: ["Is there a dog in the image?","How many people are visible?"])res["answers"].eachdo |a|
casea["verdict"]when"yes","no"thenhandle(a["verdict"])when"uncertain"thenflag_for_review(a)# the image does not settle it — a real answerwhen"n/a"thenputsa["answer"]# it wasn't a yes/no questionendendSynchronous requests are killed at 60 seconds with a 504 sync_timeout. Anything that
might run longer — a long PDF, detail: "high", a batch — belongs on the queue.
# Submit, then poll. wait_for_task handles the loop and the failure case.task=vision.analyze_and_wait(file: "contract-80-pages.pdf",preset: "contract",pages: "1-50",poll_interval: 2,max_wait: 900,on_poll: ->(t){Rails.logger.info(t["status"])})# Or submit and walk away — the result comes to you.ref=vision.analyze_async(file: "contract.pdf",preset: "contract",webhook_url: "https://yourapp.com/hooks/vision")Results stay retrievable for 7 days; after that get_task raises ResultExpiredError
(metadata survives, the payload does not).
Deliveries are signed. Verify over the raw bytes before parsing — a re-serialized body has different bytes and will not match.
classVisionHooksController < ApplicationControllerskip_before_action:verify_authenticity_tokendefcreateevent=VisionAPI::Webhook.verify(request.raw_post,request.headers["X-Vision-Signature"],ENV.fetch("VISION_WEBHOOK_SECRET"))VisionResultJob.perform_later(event)# any 2xx is success — ack fast, work afterwardshead:acceptedrescueVisionAPI::WebhookSignatureErrorhead:bad_request# never parse an unverified bodyendendVisionAPI::Webhook.verify rejects a bad signature, a malformed header and a timestamp
more than 5 minutes old, and accepts a delivery if anyv1= part matches — which is
what makes a secret rotation seamless. Get the secret from
https://app.visionapi.io/dashboard/webhooks. Failed deliveries retry at +1 m, +5 m,
+15 m and +40 m, then stop.
Every failure raises a subclass of VisionAPI::APIError carrying the HTTP status, the
stable code, and whatever details the endpoint attached. Rescue the class you mean, or
switch on code — never on the message text, which is prose and changes.
beginres=vision.analyze(file: "scan.pdf",preset: "invoice")rescueVisionAPI::InsufficientCreditsError=>ealert_ops("needs #{e.required}, has #{e.available}")# never retried — it cannot succeedrescueVisionAPI::SyncTimeoutErrortask=vision.analyze_and_wait(file: "scan.pdf",preset: "invoice")rescueVisionAPI::UnsupportedTypeErrorquarantine("not an image or a PDF")rescueVisionAPI::APIError=>eRails.logger.error("vision #{e.code} (#{e.status}) request_id=#{e.request_id}")end| Class | HTTP | Codes |
|---|---|---|
InvalidRequestError | 400 | invalid_request |
AuthenticationError | 401 | invalid_api_key, unauthorized |
InsufficientCreditsError | 402 | insufficient_credits — with #required / #available |
PermissionDeniedError | 403 | forbidden, email_not_verified |
NotFoundError | 404 | task_not_found, schema_not_found |
ConflictError | 409 | conflict |
ResultExpiredError | 410 | result_expired |
PayloadTooLargeError | 413 | file_too_large, page_limit_exceeded |
UnsupportedTypeError | 415 | unsupported_type |
UnprocessableError | 422 | pdf_encrypted, invalid_page_selection, invalid_schema, schema_field_conflict, too_many_questions |
RateLimitError | 429 | rate_limited — with #retry_after |
TooManyTasksError | 429 | too_many_tasks — the per-plan async concurrency cap, with #max_tasks. Subclasses RateLimitError, but is not auto-retried: it clears when one of your tasks finishes |
InternalError | 500 | internal_error — with #request_id |
ProviderError | 502 | provider_error |
SyncTimeoutError | 504 | sync_timeout |
UsageError (bad arguments), ConnectionError / TimeoutError (the request never got a
response) and TaskFailedError / TaskTimeoutError come from the client itself.
The client retries 429, 500, 502 and network failures — three attempts by default, with the
server's own Retry-After honored on 429 and exponential backoff with jitter elsewhere.
Input errors and insufficient_credits are never retried, because they cannot succeed.
Every billable POST is sent with a generated Idempotency-Key, so a retried upload replays
the first response instead of paying twice. Supply your own when the caller may retry — a
Sidekiq job that re-runs, a queue that redelivers — because a fresh process generates a
fresh key:
vision.analyze(file: path,preset: "invoice",idempotency_key: "invoice-#{invoice.id}")Reusing a key with a different payload raises ConflictError, which is the mechanism
working: it means the key already stands for something else.
vision=VisionAPI.new(api_key: ENV["VISION_API_KEY"],# default: ENV["VISION_API_KEY"]base_url: "https://api.visionapi.io",# default; override for a self-hosted deploymenttimeout: 120,# per request, secondsmax_retries: 3,auto_idempotency: true,headers: {"X-Trace-Id"=>trace_id}# sent on every request)Every method takes per-call idempotency_key: and timeout:.
credits=vision.creditscredits["balance"]# buckets are spent in order: subscription → rollover → pack → welcomevision.each_request(limit: 100)do |record|
puts[record["created_at"],record["endpoint"],record["preset"],record["credits_used"]].join(" ")end# each_request without a block returns an Enumerator, so this fetches one page:vision.each_request.first(10)Usage history is metadata only — never the file, never the extracted values. Uploaded files are never retained: a synchronous request holds yours in memory for the length of the call, and an async request stages it only until the worker finishes with it.
Same for everyone:
| Limit | Value |
|---|---|
| Max file size | 20 MB |
| Max PDF pages per request | 50 |
| Sync request timeout | 60 s |
Per plan:
| Limit | Free | Starter | Growth | Pro | Scale |
|---|---|---|---|---|---|
| Requests per minute, per key | 10 | 60 | 120 | 300 | 600 |
| Burst capacity | 20 | 120 | 240 | 600 | 1,200 |
| Concurrent async tasks | 1 | 4 | 8 | 16 | 32 |
| Active API keys per account | 1 | 5 | 10 | 20 | 50 |
| Saved schemas | 3 | 10 | 25 | 100 | unlimited |
Max questions per ask | 5 | 5 | 5 | 10 | 10 |
The rate-limit bucket is per API key, not per account — splitting a workload across
keys splits the limit too. The concurrency cap is per account and does not split that way:
over it, an async submission answers 429 too_many_tasks and is charged nothing.
Higher limits on paid plans: https://visionapi.io/pricing.
Runnable scripts in examples/:
| File | What it shows |
|---|---|
analyze.rb | The smallest useful call, and how to read the result |
custom_schema.rb | Custom fields, line-item injection, saved schemas |
detect_then_analyze.rb | Routing a mixed inbox before spending on extraction |
async_batch.rb | A folder of long PDFs, queued with bounded concurrency |
webhook_server.rb | A verified receiver, with no framework |
ask.rb | Visual Q&A and the verdict field |
export VISION_API_KEY=sk_live_…
ruby examples/analyze.rb invoice.pdfbundle install
rake test# offline: a stub server stands in for the API, no key needed
rubocopIssues and pull requests are welcome at https://github.com/devrobotlabs/visionapi-ruby. For anything about the API itself — a preset, a limit, an error code — https://support.visionapi.io reaches the team faster.
MIT © Vision API