EXPLOREMeet your next model. Discover ArtRuby Model APIs
Model APIs

Your next great idea.
Meet its intelligence.

Explore Qwen, GLM, Kimi and DeepSeek. Compare per-token rates and choose a model family for your application, without starting from bare metal.

Models, without the guesswork.

USD / 1M tokens · Final rates · Cached input priced separately

Swipe the table to compare input, output and cache rates →

Small models

1 model
Small models API prices in US dollars per million tokens
ModelInputOutputCached inputGet started
Qwen3.8-27BQwen family
$0.21$1.28$0.043Request access

Large models

4 models
Large models API prices in US dollars per million tokens
ModelInputOutputCached inputGet started
GLM-5.3GLM family
$0.62$2.16$0.16Request access
GLM-5.3-FlashGLM family
$0.066$0.23$0.019Request access
Kimi-K3Kimi family
$1.95$9.75$0.20Request access
DeepSeek-V4.1-FlashDeepSeek family
$0.081$0.32$0.006Request access

5 models in the catalog

A familiar starting point

From an idea
to your first request.

Get the integration details you need from our team, then connect your application using the assigned endpoint and model identifier.

first-request.sh
# Use the endpoint & model ID from onboarding
curl "$ARTRUBY_BASE_URL/chat/completions" \
  -H "Authorization: Bearer $ARTRUBY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "YOUR_ASSIGNED_MODEL_ID",
    "messages": [{
      "role": "user",
      "content": "Let’s build something great."
    }]
  }' 
Illustrative requestcurl / JSON

Illustrative integration only. Access is confirmed during onboarding; this site does not issue API keys or execute model requests.

A few things to know

Good questions.
Clear answers.

Have something else in mind?
Talk to our team ↗

How do I get API access?

Choose a model and select “Request access”. Share your expected traffic, use case and deployment requirements. The team will confirm availability, credentials, the base URL and exact model identifier.

How does cached-input pricing work?

Eligible cache-hit input tokens use the cached-input rate instead of the standard input rate. Cache eligibility and usage reporting are determined by the serving configuration, not by the calculator on this website.

What happens to my prompts and completions?

Please review our Data Policy for the existing inference-content and operational-metadata policies, and confirm the scope for your specific service during onboarding.

Let’s build what’s next

The next thing you build
starts with the right foundation.

Tell us what you’re working on. We’ll help you find the right fit.