EXPLOREMeet your next model. Discover ArtRuby Model APIs
Transparent by design

Ambitious ideas.
Grounded pricing.

Compare GPU plans and model rates in US dollars. No mystery math. Just a clear starting point for your next build.

Models, without the guesswork.

USD / 1M tokens · Final rates · Cached input priced separately

Swipe the table to compare input, output and cache rates →

Model API prices in US dollars per million tokens
ModelInputOutputCached inputGet started
Qwen3.8-27BQwen family
$0.2150% off benchmark$1.2850% off benchmark$0.04350% off benchmarkRequest access
GLM-5.3GLM family
$0.6256% off benchmark$2.1651% off benchmark$0.1640% off benchmarkRequest access
GLM-5.3-FlashGLM family
$0.06656% off benchmark$0.2354% off benchmark$0.01937% off benchmarkRequest access
Kimi-K3Kimi family
$1.9535% off benchmark$9.7535% off benchmark$0.2035% off benchmarkRequest access
DeepSeek-V4.1-FlashDeepSeek family
$0.08173% off benchmark$0.3273% off benchmark$0.006Benchmark rateRequest access
Clear rates. No second discount.Rates include the reductions shown against our recorded model-developer benchmarks. Discounts vary by token category. Cached input replaces the standard input rate for eligible cache hits.

Benchmarks use the per-model reference dates listed in the pricing details, from official provider listings on OpenRouter. Compact prices may be rounded for display; the calculator uses exact rates. View exact rates and pricing basis ↗

Find your compute.

One GPU per container. Memory and CPU included as listed.

Swipe the table to compare all specifications and prices →

GPU cloud container prices in US dollars
GPU modelVRAMSystem RAMCPU allocationPrice / hourGet started
NVIDIA RTX 3060Cloud container
12 GB32 GB14 cores$0.13/hrRequest capacity
NVIDIA RTX 3090Cloud container
24 GB32 GB16 cores$0.21/hrRequest capacity
NVIDIA A100 PCIeCloud container
40 GB64 GB10 cores$0.51/hrRequest capacity
NVIDIA A800 PCIeCloud container
80 GB100 GB14 cores$1.05/hrRequest capacity
NVIDIA H800 PCIeCloud container
80 GB100 GB20 cores$2.39/hrRequest capacity
NVIDIA H100 NVLinkCloud container
80 GB200 GB20 cores$2.54/hrRequest capacity
A plan for your workload.Hourly and monthly rates are separate plans, not prorated equivalents. Region, stock, storage, bandwidth and billing terms are confirmed with your quote.
Built around your needs

A configuration, not a compromise.

Cloud infrastructure is quoted to match your capacity, location and network requirements.

Bare metal

Dedicated hardware, chosen for the way you work. Specify the compute, memory and storage your applications need.

Content delivery

Bring your content closer to your audience. Plan delivery for websites, applications and media with a network setup tailored to your traffic.

DDoS protection

Build resilience into your network plan. Discuss protection for your internet-facing applications and infrastructure.

Cloud storage

Give your data a dependable foundation. Match capacity, access patterns and deployment requirements to the way your business grows.

Prices are in US dollars. Final pricing, taxes and availability are confirmed before purchase. GPU prices are per cloud container. API prices are per million tokens.

Plan before you build

A little clarity goes a long way.

Estimate the cost of your model usage. Adjust tokens and cache hits to see how the numbers add up.

Share of total input tokens.
Estimated usage cost · USD$1.12
Standard input$0.66
Cached input$0.00
Output$0.46

Calculated using unrounded discounted rates. An estimate, not a quote. Excludes taxes and any separately agreed charges.

A price you can trace

A clear basis. A fixed rate.

Our API rates use a recorded USD benchmark from each model developer’s official provider listing on OpenRouter, not another provider’s promotional price. Each token category has its own reduction. These are fixed ArtRuby rates, not a live market quote.

View exact rates and reference prices

USD per million tokens. Values are ordered input / output / cached input. “Pay” is the share of the benchmark charged; for example, pay 49% means 51% off. A 100% factor means no discount.

API benchmark rates, payment factors and exact final USD rates
Model / official providerRecorded benchmarkPay (% of benchmark)Exact ArtRuby rate
Qwen3.8-27B ↗Alibaba · 2026-09-24$0.425 / $2.55 / $0.08550% / 50% / 50%$0.2125 / $1.275 / $0.0425
GLM-5.3 ↗Z.AI · 2026-09-22$1.4 / $4.4 / $0.2644% / 49% / 60%$0.616 / $2.156 / $0.156
GLM-5.3-Flash ↗Z.AI · 2026-09-22$0.15 / $0.5 / $0.0344% / 46% / 63%$0.066 / $0.23 / $0.0189
Kimi-K3 ↗Moonshot AI · 2026-09-22$3 / $15 / $0.365% / 65% / 65%$1.95 / $9.75 / $0.195
DeepSeek-V4.1-Flash ↗DeepSeek · 2026-09-22$0.3 / $1.2 / $0.00627% / 27% / 100%$0.081 / $0.324 / $0.006

Each reference date is shown with its official provider. Cache creation, storage and separate cache products are not included in this comparison. Model names and benchmark links do not imply a partnership or endorsement.

A few things to know

Good questions.
Clear answers.

Have something else in mind?
Talk to our team ↗

Have the model discounts already been applied?

Yes. Rates already include the reductions shown against our recorded official-provider benchmarks. Input, output and cached input can have different reductions. DeepSeek cached input uses the reference rate without an extra discount. No further discount should be applied.

How are prices displayed?

API prices use a compact display format, with extra precision for smaller rates. The cost calculator uses the full underlying rates and rounds only the final dollar estimate.

Can I reserve capacity or discuss larger volumes?

Contact our team with your model, GPU configuration, region, expected volume and timing. Availability and any custom commercial terms will be confirmed in writing.

Let’s build what’s next

The next thing you build
starts with the right foundation.

Tell us what you’re working on. We’ll help you find the right fit.