Skip to content

Repository files navigation

Vision API — Ruby client

Official Ruby client for Vision API — send an image or a PDF, describe the fields you want in plain language, get structured JSON back with a confidence level on every value.

Gem Versionlicense


Install

# Gemfilegem"vision_api"
gem install vision_api

Ruby 3.0+. No runtime dependencies — net/http, json and openssl are all standard library.

Quick start

require"vision_api"vision=VisionAPI.new# reads ENV["VISION_API_KEY"]res=vision.analyze(file: "invoice.pdf",preset: "invoice")res["result"]["invoice_id"]["value"]# => "A-10422"res["result"]["total"]["value"]# => 1284.5, or nil if the invoice has no totalres["credits_used"]# => 3

Requests are metered in credits, per image and per selected PDF page — see pricing for current rates. Failures cost nothing: the reservation is released in full on any non-2xx, so there is no compensating logic to write.

Server-side only. There is no publishable key and no test mode — an API key is a live spending credential. Keep it in the environment (or Rails credentials), never in a repository and never in anything a browser downloads.


Reading a result

Responses are plain hashes with the wire's keys, so everything you already know about hashes applies. Two rules explain almost every surprise:

1. Every scalar is wrapped.{"value" => …, "confidence" => "low"|"mid"|"high"}. Read res["result"]["total"]["value"], not res["result"]["total"].

2. A preset response contains every field of that preset — including the ones the document does not carry, which come back as {"value" => nil, "confidence" => "low"}. A key being present does not mean a value was found.

Line-item arrays are the one shape worth looking at twice. The array itself is not wrapped; each cell inside each row is:

{"invoice_id"=>{"value"=>"A-10422","confidence"=>"high"},"carrier"=>{"value"=>nil,"confidence"=>"low"},"line_item"=>[{"description"=>{"value"=>"Widget","confidence"=>"high"},"quantity"=>{"value"=>2,"confidence"=>"high"},"amount"=>{"value"=>25.0,"confidence"=>"mid"}}]}

VisionAPI::Result covers the common readings, so you rarely have to spell that out:

includeVisionAPI::Result# or call them as VisionAPI::Result.unwrap(…)unwrap(res["result"])# => {"invoice_id" => "A-10422", "carrier" => nil, "line_item" => [{"description" => "Widget", …}]}unwrap(res["result"],drop_null: true)# only what was actually foundvalue(res["result"],"total",0)# 1284.5, or 0 when absentrows(res["result"],"line_item")# [] when the invoice has no linespresent(res["result"])# ["invoice_id", "total", "line_item"]missing(res["result"])# ["carrier", …]below_confidence(res["result"],"high")# fields to route to a human

What you can send

Exactly one file source per call:

vision.analyze(file: "invoice.pdf",preset: "invoice")# a pathvision.analyze(file: Pathname("invoice.pdf"),preset: "invoice")# a Pathnamevision.analyze(file: File.open("invoice.pdf","rb"),)# an open binary IOvision.analyze(file: ["scan.png",bytes],)# bytes + a namevision.analyze(file_url: "https://example.com/invoice.pdf",)# a public URLvision.analyze(file_base64: encoded,)# "data:" prefix optional

JPEG, PNG, WebP, TIFF and PDF, up to 20 MB and 50 pages. The type is detected from magic bytes — the filename is ignored.

Options

KeywordDefaultWhat it does
preset:A catalog name, or "auto" to let the API classify the file first (free).
schema:Custom fields, alone or on top of a preset.
schema_name:A schema saved in your dashboard. Excludes preset: and schema:.
pages:allPDF page selection, e.g. "1-3,7". You pay for selected pages only.
language_hint:autoISO 639-1 code, e.g. "es".
detail:"standard""high" renders pages at higher resolution. Same cost, slower.
output:"json""text" returns raw OCR text instead of fields.
include_raw_text:falseAdds full_text, the whole transcription, alongside result.
min_confidence:"low"Fields below the level come back nil, with confidence preserved.

Custom fields

A schema is a flat hash: each key is a field name, each value describes what to extract. It is compiled before any credit moves, so a bad schema costs nothing.

res=vision.analyze(file: "invoice.pdf",preset: "invoice",schema: {# Plain form — the string is the description, type defaults to string."machine_serial"=>'Serial number of the machine being invoiced, without the "SN:" prefix',# Typed form."total_net"=>{"type"=>"number","description"=>"Total before tax"},"signed_on"=>{"type"=>"date","description"=>"Date the contract was signed"},# Reserved key: injects fields into every row of the preset's line-item array."line_item"=>{"lot_number"=>"The lot number printed on the line, if present"}})

Field names must match ^[a-z][a-z0-9_]{0,63}$. Types are string (default), number, boolean, date, array and object. A custom name that collides with a preset field is a 422 schema_field_conflict — rename it, or use the preset's own field.

Descriptions are the prompt. "The invoice number exactly as printed, without the #" extracts better than "invoice number". Say what to do when the value is missing or ambiguous if it matters.

Reuse a combination by saving it:

vision.create_schema("our-invoices",preset: "invoice",schema: {"machine_serial"=>"…"})vision.analyze(file: "invoice.pdf",schema_name: "our-invoices")

Picking a preset

28 presets ship with the API. Fetch the catalog rather than hardcoding field names from memory — presets are versioned, and the catalog is the source of truth:

vision.presets.each{ |p| puts"#{p['name']} (#{p['kind']}) — #{p['field_count']} fields"}vision.preset("invoice")["fields"].map{ |f| f["name"]}

Three ways to choose:

# 1. You know what it is.vision.analyze(file: "receipt.jpg",preset: "receipt")# 2. You don't, and you want the data anyway. Classification is free.res=vision.analyze(file: "unknown.pdf",preset: "auto")res["detection"]["preset"]# what ranres["detection"]["fallback"]# true = "shape unknown", not a matchres["detection"]["alternatives"]# the rest of the ranking, best first# 3. The *type* is the decision — routing a mixed inbox, or refusing to spend# on a 40-page PDF until you know what it is. Far cheaper than extracting.guess=vision.detect(file: "unknown.pdf")vision.analyze(file: "unknown.pdf",preset: guess["recommended"])unlessguess["fallback"]

detect reads page 1 only, so an image and a 300-page PDF cost the same, and it is metered in batches rather than per call: most calls report credits_used 0 and an occasional one carries the charge. See pricing for the rate.


Questions instead of fields

Up to 5 questions about one file, priced exactly like an extraction. The questions themselves are free.

res=vision.ask(file: "photo.jpg",questions: ["Is there a dog in the image?","How many people are visible?"])res["answers"].eachdo |a|
casea["verdict"]when"yes","no"thenhandle(a["verdict"])when"uncertain"thenflag_for_review(a)# the image does not settle it — a real answerwhen"n/a"thenputsa["answer"]# it wasn't a yes/no questionendend

Long jobs: async and webhooks

Synchronous requests are killed at 60 seconds with a 504 sync_timeout. Anything that might run longer — a long PDF, detail: "high", a batch — belongs on the queue.

# Submit, then poll. wait_for_task handles the loop and the failure case.task=vision.analyze_and_wait(file: "contract-80-pages.pdf",preset: "contract",pages: "1-50",poll_interval: 2,max_wait: 900,on_poll: ->(t){Rails.logger.info(t["status"])})# Or submit and walk away — the result comes to you.ref=vision.analyze_async(file: "contract.pdf",preset: "contract",webhook_url: "https://yourapp.com/hooks/vision")

Results stay retrievable for 7 days; after that get_task raises ResultExpiredError (metadata survives, the payload does not).

Verifying a delivery

Deliveries are signed. Verify over the raw bytes before parsing — a re-serialized body has different bytes and will not match.

classVisionHooksController < ApplicationControllerskip_before_action:verify_authenticity_tokendefcreateevent=VisionAPI::Webhook.verify(request.raw_post,request.headers["X-Vision-Signature"],ENV.fetch("VISION_WEBHOOK_SECRET"))VisionResultJob.perform_later(event)# any 2xx is success — ack fast, work afterwardshead:acceptedrescueVisionAPI::WebhookSignatureErrorhead:bad_request# never parse an unverified bodyendend

VisionAPI::Webhook.verify rejects a bad signature, a malformed header and a timestamp more than 5 minutes old, and accepts a delivery if anyv1= part matches — which is what makes a secret rotation seamless. Get the secret from https://app.visionapi.io/dashboard/webhooks. Failed deliveries retry at +1 m, +5 m, +15 m and +40 m, then stop.


Errors

Every failure raises a subclass of VisionAPI::APIError carrying the HTTP status, the stable code, and whatever details the endpoint attached. Rescue the class you mean, or switch on code — never on the message text, which is prose and changes.

beginres=vision.analyze(file: "scan.pdf",preset: "invoice")rescueVisionAPI::InsufficientCreditsError=>ealert_ops("needs #{e.required}, has #{e.available}")# never retried — it cannot succeedrescueVisionAPI::SyncTimeoutErrortask=vision.analyze_and_wait(file: "scan.pdf",preset: "invoice")rescueVisionAPI::UnsupportedTypeErrorquarantine("not an image or a PDF")rescueVisionAPI::APIError=>eRails.logger.error("vision #{e.code} (#{e.status}) request_id=#{e.request_id}")end
ClassHTTPCodes
InvalidRequestError400invalid_request
AuthenticationError401invalid_api_key, unauthorized
InsufficientCreditsError402insufficient_credits — with #required / #available
PermissionDeniedError403forbidden, email_not_verified
NotFoundError404task_not_found, schema_not_found
ConflictError409conflict
ResultExpiredError410result_expired
PayloadTooLargeError413file_too_large, page_limit_exceeded
UnsupportedTypeError415unsupported_type
UnprocessableError422pdf_encrypted, invalid_page_selection, invalid_schema, schema_field_conflict, too_many_questions
RateLimitError429rate_limited — with #retry_after
TooManyTasksError429too_many_tasks — the per-plan async concurrency cap, with #max_tasks. Subclasses RateLimitError, but is not auto-retried: it clears when one of your tasks finishes
InternalError500internal_error — with #request_id
ProviderError502provider_error
SyncTimeoutError504sync_timeout

UsageError (bad arguments), ConnectionError / TimeoutError (the request never got a response) and TaskFailedError / TaskTimeoutError come from the client itself.

Retries and idempotency

The client retries 429, 500, 502 and network failures — three attempts by default, with the server's own Retry-After honored on 429 and exponential backoff with jitter elsewhere. Input errors and insufficient_credits are never retried, because they cannot succeed.

Every billable POST is sent with a generated Idempotency-Key, so a retried upload replays the first response instead of paying twice. Supply your own when the caller may retry — a Sidekiq job that re-runs, a queue that redelivers — because a fresh process generates a fresh key:

vision.analyze(file: path,preset: "invoice",idempotency_key: "invoice-#{invoice.id}")

Reusing a key with a different payload raises ConflictError, which is the mechanism working: it means the key already stands for something else.


Configuration

vision=VisionAPI.new(api_key: ENV["VISION_API_KEY"],# default: ENV["VISION_API_KEY"]base_url: "https://api.visionapi.io",# default; override for a self-hosted deploymenttimeout: 120,# per request, secondsmax_retries: 3,auto_idempotency: true,headers: {"X-Trace-Id"=>trace_id}# sent on every request)

Every method takes per-call idempotency_key: and timeout:.


Account and usage

credits=vision.creditscredits["balance"]# buckets are spent in order: subscription → rollover → pack → welcomevision.each_request(limit: 100)do |record|
puts[record["created_at"],record["endpoint"],record["preset"],record["credits_used"]].join(" ")end# each_request without a block returns an Enumerator, so this fetches one page:vision.each_request.first(10)

Usage history is metadata only — never the file, never the extracted values. Uploaded files are never retained: a synchronous request holds yours in memory for the length of the call, and an async request stages it only until the worker finishes with it.


Limits

Same for everyone:

LimitValue
Max file size20 MB
Max PDF pages per request50
Sync request timeout60 s

Per plan:

LimitFreeStarterGrowthProScale
Requests per minute, per key1060120300600
Burst capacity201202406001,200
Concurrent async tasks1481632
Active API keys per account15102050
Saved schemas31025100unlimited
Max questions per ask5551010

The rate-limit bucket is per API key, not per account — splitting a workload across keys splits the limit too. The concurrency cap is per account and does not split that way: over it, an async submission answers 429 too_many_tasks and is charged nothing. Higher limits on paid plans: https://visionapi.io/pricing.


Examples

Runnable scripts in examples/:

FileWhat it shows
analyze.rbThe smallest useful call, and how to read the result
custom_schema.rbCustom fields, line-item injection, saved schemas
detect_then_analyze.rbRouting a mixed inbox before spending on extraction
async_batch.rbA folder of long PDFs, queued with bounded concurrency
webhook_server.rbA verified receiver, with no framework
ask.rbVisual Q&A and the verdict field
export VISION_API_KEY=sk_live_…
ruby examples/analyze.rb invoice.pdf

Development

bundle install
rake test# offline: a stub server stands in for the API, no key needed
rubocop

Contributing

Issues and pull requests are welcome at https://github.com/devrobotlabs/visionapi-ruby. For anything about the API itself — a preset, a limit, an error code — https://support.visionapi.io reaches the team faster.

License

MIT © Vision API

About

Official Ruby client for the Vision API. Extract structured JSON from images and PDFs, with a confidence level on every field.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages