Blog
How Ofrai routes prompts across models
Par Ofrai Editorial (Editorial Team)
A prompt router is the layer between the user's text box and a forest of upstream models — chat, image, video, and audio. On Ofrai every request passes through one router before it touches a vendor API. We built it because no single model wins every job, and we did not want our users to keep three subscriptions open to find out which one does. This is a field note on what the router actually does, what it costs us to run, where it breaks, and what we are shipping next.
Why we built a router
Take the same prompt — 'a glass jar of coffee beans on a wood table, soft morning light' — and run it across five live engines. On Ofrai that prompt goes to Nano Banana Pro in about 8 seconds for 220 credits, to GPT Image 2 in about 12 seconds for 320 credits, to Flux Kontext in about 6 seconds for 180 credits, to Seedream 5 in roughly 9 seconds for 260 credits, and to the legacy GPT Image 1.5 in about 4 seconds for 80 credits. Same prompt, five wallets. Outside Ofrai that is five browser tabs, five logins, and a CSV at the end of the day. Inside Ofrai it is one click in Image Studio and a swap from the model dropdown. The router's job is to make that swap invisible — to pick the cheapest engine that meets a quality bar, and to keep the user from having to learn a vendor's API quirks to know which engine is which. The honest version is that even our model leads disagree on which image engine is best for a given brief. Nano Banana Pro nails photographic lighting; GPT Image 2 nails text-in-image; Flux Kontext is the fastest. None of them wins on every dimension, and on some prompts the cheapest engine wins outright. Routing is the honest acknowledgment that model quality is a vector, not a scalar.
Our routing rules
The router's input is the prompt plus a small bag of context — user locale, credit balance, previous attempts, the studio the prompt came from. The output is a vendor, a model id, and a timeout. Five rules do most of the work. First, studio wins. If the user is in Video Studio, the router only considers video engines — Seedance 2.5, Seedance 2, Kling 3, Sora 2, Veo 3, Wan 2.6. Cross-kind routing is opt-in only. Second, prompt shape. Short prompts under 80 characters default to the cheapest engine that meets our quality floor; long, structured prompts with weight hints default to the highest-quality engine the wallet can afford. Third, language. We detect language on the prompt before routing chat calls. Korean and Japanese traffic prefers the multilingual chat models we keep warm; English routes by cost first. Fourth, credit fallback. If the preferred engine costs more credits than the user has left, we drop one rung down the ladder, not all the way. A user with 100 credits who picks a 320-credit job still gets the next-best 220-credit engine, not a flat error. Fifth, retry chain. On 5xx or upstream timeout we retry the same engine twice with exponential backoff (250 ms then 750 ms), then fall over to a peer engine in the same studio. We do not retry moderation rejections. The router is a single TypeScript module called routePrompt, not a microservice. It runs inside the Next.js route handler, so it inherits the same auth, rate limit, and tracing context as the studio it serves. That decision cost us a few weekends during the streaming rewrite and saved us a deploy.
Failure modes we ship around
We did not anticipate most of these. Sharing because they are common to anyone building a multi-vendor stack. Vendor 5xx storms are routine — one provider had a 40-minute partial outage on a Tuesday afternoon and our retry logic ate it cleanly because each studio had a peer engine warm; if you only integrate one vendor per studio, you eat the outage. Moderation drift is sneakier — two vendors tightened their safety classifiers inside a week and started rejecting prompts that had shipped the previous month. The router does not adjudicate moderation, it surfaces the rejection, but we now cache the last-known-good prompt hash per user so we can show 'this used to work, here is the rejection reason.' Credit ledger drift is the accounting tax of multi-vendor stacks: vendors occasionally bill less than our catalog says (a 320-credit job bills 310, a 6000-credit job bills 5800), and we log the delta and refund it nightly; the bigger problem is the reverse, where vendors sometimes bill more than they delivered, and we alert on a 5 percent drift and freeze the route until support replies. Timeout for video was the simplest fix — our initial 60-second ceiling was wrong for video, so we shipped 90 seconds and added a heartbeat poll so the UI can show 'rendering frame 12 of 90' instead of a silent spinner. Image seed collisions were the most embarrassing: two users hitting the same engine at the same time with near-identical prompts got identical outputs, and we now stamp a job-local seed onto every prompt before it leaves Ofrai.
How we measure
Three dashboards, refreshed hourly. The first is latency — p50 and p95 per studio per engine, in seconds. Chat should be under 12 seconds p95; image under 30; video under 90 for a 5-second clip. Anything trending above the line for an hour pages the on-call. The second is cost per completion — credits consumed per successful generation, broken out by engine. We watch this for regressions: if Seedance 2.5 p95 cost-per-job jumps 20 percent week over week, the engine's output behavior changed and we re-evaluate. The third is a quality spot-check — fifty prompts per week, hand-graded by Yuki's team, run across every live engine. The grades are private; the deltas are public. We do not measure user satisfaction directly. We measure what users do next: did they regenerate with the same engine, did they switch engines, did they download the output, did they delete it. Re-generate-without-switch is our quietest signal that the router picked right. We also keep a fallback counter per studio — how often a peer engine rescued a job — because a healthy fallback rate means the retry chain is doing its job and a zero fallback rate over a week usually means an upstream is silently degraded.
What's next
Three things, in order. First, a cache layer keyed on a stable hash of the prompt plus the user locale plus the resolved engine — same prompt and locale inside a 24-hour window should reuse the last response unless the user opts out. The privacy question is harder than the caching question; we will surface the toggle in the studio. Second, streaming responses end-to-end so the chat studio shows tokens as they arrive, which means the router has to resolve the engine before the first token, not before the request. Third, a user feedback loop — thumbs-up, thumbs-down on the model card — so the router can learn from a user's studio history instead of a global quality floor. None of this is shipped yet. All of it is in progress. If you are building your own router, the only advice we would give is to start from the studio, not from the model. Studios know what the user is trying to make. Models know how to make it. The router's job is to translate.
FAQ
- What is a prompt router?
- A prompt router is the layer between your input and a group of AI models. It reads the prompt, the user context, and the live state of each model, then picks one engine to send the request to. The goal is to match the right model to each job instead of forcing one model to handle everything.
- How does prompt routing work?
- On Ofrai the router inspects the prompt length, language, the studio the user is in (chat, image, or video), and the user's credit balance, then resolves to a single vendor and model id. If that engine fails with a 5xx or times out, the router retries the same engine twice with exponential backoff before falling over to a peer engine in the same studio.
- How does Ofrai choose which AI model to use?
- The router follows five rules. Studio wins (a Video Studio prompt only sees video engines), short prompts default to the cheapest engine that meets quality, language detection steers chat calls to the multilingual models, credit balance triggers a one-rung downgrade instead of a hard error, and failures retry the same engine twice before falling over to a peer.
- Does routing increase latency?
- Routing itself adds under 20 ms in Ofrai — it is a synchronous TypeScript module, not a network hop. The latency you see is almost entirely upstream render time: chat 2–12 s, image 8–30 s, video 30–90 s per 5-second clip. Picking the right engine usually cuts wall-clock time because the router avoids the slowest peer.
- What is a model router?
- A model router is the routing logic that decides which AI model handles a given request. In a single-model product the router is trivial — there is nothing to choose. In a multi-model product like Ofrai it has to weigh cost, quality, latency, and current upstream health, then commit to one engine and own the retry path when that engine fails.