Save on every task. Tailored to your needs. With Gemini, Claude, OpenAI, DeepSeek, Qwen, Meta, Mistral, GLM, MiniMax, Kimi, NVIDIA

01 / Product

Hydra selects a suitable model for each request. Context, complexity and your model pool guide the choice.

We evaluate every model across different tasks to understand its strengths and limitations. The results help Hydra select a suitable model for your request.

Hydra weighs quality, expected response time and cost. The API returns the model selection; your application calls the selected model.

Model selection25+ models
Your requestContext · complexity
hydra
Model XModel YModel Z
Illustrative model selection

Hosted in Germany

The routing engine runs in Germany.

25+ models

One routing pool for different kinds of tasks.

Always up to date.

We add evaluated new models. Your routing stays up to date automatically.

Your own model pool

Choose which models Hydra may consider for your requests. Hydra selects a suitable model from your pool.

Open-source filter

Optionally limit selection to models classified as open source in the Hydra pool.

Pseudonymization

Optional and self-hosted: detected personal data is masked before routing and model execution.

Native Hydra API

02 / Results

High quality.
Lower costs.

Lower costs. More room to build.

≈89%

lower generation cost in our internal benchmark

Hydra D vs. GPT-5.6 Luna max · 100 planned tasks

A broad range of text tasks across different topics and complexity levels, including reasoning. Measurements cover direct responses without tool use.

Excludes coordinator, judging and operations. Historical comparison, not a savings guarantee.

Mean quality

Points / 100

Historical Hydra configurationsReference models
GLM 5.3 Flash99/100 scored
87.95 / 100; 95% interval: 82.78–92.70
DeepSeek V4 Pro 081399/100 scored
87.72 / 100; 95% interval: 82.31–92.73
Hydra D98/100 scored
85.43 / 100; 95% interval: 79.71–90.55
GPT-5.6 Luna max99/100 scored
84.43 / 100; 95% interval: 78.54–89.94
Hydra A99/100 scored
84.06 / 100; 95% interval: 77.87–90.00
MiniMax M399/100 scored
83.91 / 100; 95% interval: 78.32–89.18
Hydra H95/100 scored
83.83 / 100; 95% interval: 77.50–89.89
GPT-OSS 120B high95/100 scored
82.82 / 100; 95% interval: 76.34–88.37
GPT-5.6 Luna low99/100 scored
81.77 / 100; 95% interval: 75.55–87.73

Mean over scored tasks. Lines: 95% bootstrap intervals. Between 95 and 99 of 100 tasks were scored per configuration.

Internal measurement · no independent holdout. Hydra A, D and H are historical configurations, not results for today’s Performance and Balanced profiles.

Methodology & full data

100 planned text tasks per configuration, direct responses without tools. Scored with Fable 5.1; means include scored tasks only. The 95% intervals use 2,000 task-bootstrap samples and are descriptive. The development set was reused, and measurement waves had different conditions. This does not establish a general ranking at equivalent quality.

Costs include recorded generation attempts for 100 planned tasks. Unresolved attempt costs make the marked values lower bounds. Coordinator, judging and operations are excluded. Hydra H was measured with an open-source filter.

Source: technical white paper, 11 Sep 2026, table 1, pp. 5–6 ↗

Nine selected historical configurations · internal benchmark
ConfigurationQuality / 10095% intervalScoredUSD / 100
GLM 5.3 Flash87.9582.78–92.7099/100$0.11406
DeepSeek V4 Pro 081387.7282.31–92.7399/100$0.39692
Hydra D85.4379.71–90.5598/100$0.02062
GPT-5.6 Luna max84.4378.54–89.9499/100$0.18695
Hydra A84.0677.87–90.0099/100≥ $0.03989
MiniMax M383.9178.32–89.1899/100$0.20855
Hydra H · OSS83.8377.50–89.8995/100≥ $0.04837
GPT-OSS 120B high82.8276.34–88.3795/100≥ $0.14708
GPT-5.6 Luna low81.7775.55–87.7399/100$0.02974

03 / Developers

One API.
Your application.

Send your request and a profile. Hydra returns the routing decision; your application calls the selected model.

API documentation

Application example

# Your application API settings
[hydra]
endpoint = "https://hydraroute.com/v1/answers"
api_key = "YOUR_API_KEY"
profile = "performance"
open_source = false
UTF-8TOML

Example application settings. Hydra accepts JSON API requests.

Requires approved access and your own API key.

How data is processed

The routing engine runs in Germany. The console and API proxy use Cloudflare. The console stores technical request metadata, not prompt or response content. Connected providers apply their own processing terms.

Optional pseudonymization in the native Hydra API replaces detected personal data before further processing. Responses remain pseudonymized; original values are not restored automatically.

04 / ROUTING INSIGHTS

Routing in detail

Requests by model

Model X44%
Model Y36%
Model Z20%

Understand model selection

Request shares show how usage is distributed across models.

Generation costs$0.24 USD
Requests120
per request$0.002

Make sense of costs

Put generation costs in perspective, over time and per request.

Response time · median1.8 s
Last 7 daysp50 · 1.8 s
p952.9 s

Understand response times

Consider the median alongside slower responses for a fuller picture.

Illustrative view, not live data. Generation costs and response times refer to model execution; the routing API returns the model selection.

Join the list for early access to Hydra.

Request access

hydra

Hydra / The film0:57