Bare metal
Dedicated hardware, chosen for the way you work. Specify the compute, memory and storage your applications need.
Compare GPU plans and model rates in US dollars. No mystery math. Just a clear starting point for your next build.
Swipe the table to compare input, output and cache rates →
| Model | Input | Output | Cached input | Get started |
|---|---|---|---|---|
Qwen3.8-27BQwen family | $0.2150% off benchmark | $1.2850% off benchmark | $0.04350% off benchmark | Request access |
GLM-5.3GLM family | $0.6256% off benchmark | $2.1651% off benchmark | $0.1640% off benchmark | Request access |
GLM-5.3-FlashGLM family | $0.06656% off benchmark | $0.2354% off benchmark | $0.01937% off benchmark | Request access |
Kimi-K3Kimi family | $1.9535% off benchmark | $9.7535% off benchmark | $0.2035% off benchmark | Request access |
DeepSeek-V4.1-FlashDeepSeek family | $0.08173% off benchmark | $0.3273% off benchmark | $0.006Benchmark rate | Request access |
No models match your search. Try another name or select “All models”.
Benchmarks use the per-model reference dates listed in the pricing details, from official provider listings on OpenRouter. Compact prices may be rounded for display; the calculator uses exact rates. View exact rates and pricing basis ↗
Swipe the table to compare all specifications and prices →
| GPU model | VRAM | System RAM | CPU allocation | Price / hour | Get started |
|---|---|---|---|---|---|
NVIDIA RTX 3060Cloud container | 12 GB | 32 GB | 14 cores | $0.13/hr | Request capacity |
NVIDIA RTX 3090Cloud container | 24 GB | 32 GB | 16 cores | $0.21/hr | Request capacity |
NVIDIA RTX 4090Cloud container | 24 GB | 64 GB | 16 cores | $0.30/hr | Request capacity |
NVIDIA A100 PCIeCloud container | 40 GB | 64 GB | 10 cores | $0.51/hr | Request capacity |
NVIDIA A800 PCIeCloud container | 80 GB | 100 GB | 14 cores | $1.05/hr | Request capacity |
NVIDIA H800 PCIeCloud container | 80 GB | 100 GB | 20 cores | $2.39/hr | Request capacity |
NVIDIA H100 NVLinkCloud container | 80 GB | 200 GB | 20 cores | $2.54/hr | Request capacity |
Cloud infrastructure is quoted to match your capacity, location and network requirements.
Dedicated hardware, chosen for the way you work. Specify the compute, memory and storage your applications need.
Bring your content closer to your audience. Plan delivery for websites, applications and media with a network setup tailored to your traffic.
Build resilience into your network plan. Discuss protection for your internet-facing applications and infrastructure.
Give your data a dependable foundation. Match capacity, access patterns and deployment requirements to the way your business grows.
Prices are in US dollars. Final pricing, taxes and availability are confirmed before purchase. GPU prices are per cloud container. API prices are per million tokens.
Estimate the cost of your model usage. Adjust tokens and cache hits to see how the numbers add up.
Our API rates use a recorded USD benchmark from each model developer’s official provider listing on OpenRouter, not another provider’s promotional price. Each token category has its own reduction. These are fixed ArtRuby rates, not a live market quote.
USD per million tokens. Values are ordered input / output / cached input. “Pay” is the share of the benchmark charged; for example, pay 49% means 51% off. A 100% factor means no discount.
| Model / official provider | Recorded benchmark | Pay (% of benchmark) | Exact ArtRuby rate |
|---|---|---|---|
| Qwen3.8-27B ↗Alibaba · 2026-09-24 | $0.425 / $2.55 / $0.085 | 50% / 50% / 50% | $0.2125 / $1.275 / $0.0425 |
| GLM-5.3 ↗Z.AI · 2026-09-22 | $1.4 / $4.4 / $0.26 | 44% / 49% / 60% | $0.616 / $2.156 / $0.156 |
| GLM-5.3-Flash ↗Z.AI · 2026-09-22 | $0.15 / $0.5 / $0.03 | 44% / 46% / 63% | $0.066 / $0.23 / $0.0189 |
| Kimi-K3 ↗Moonshot AI · 2026-09-22 | $3 / $15 / $0.3 | 65% / 65% / 65% | $1.95 / $9.75 / $0.195 |
| DeepSeek-V4.1-Flash ↗DeepSeek · 2026-09-22 | $0.3 / $1.2 / $0.006 | 27% / 27% / 100% | $0.081 / $0.324 / $0.006 |
Each reference date is shown with its official provider. Cache creation, storage and separate cache products are not included in this comparison. Model names and benchmark links do not imply a partnership or endorsement.
Yes. Rates already include the reductions shown against our recorded official-provider benchmarks. Input, output and cached input can have different reductions. DeepSeek cached input uses the reference rate without an extra discount. No further discount should be applied.
API prices use a compact display format, with extra precision for smaller rates. The cost calculator uses the full underlying rates and rounds only the final dollar estimate.
Contact our team with your model, GPU configuration, region, expected volume and timing. Availability and any custom commercial terms will be confirmed in writing.
Tell us what you’re working on. We’ll help you find the right fit.