Free model routesPrivateHigh confidence

Z.AI Model API

3 model routes currently listed at zero price.

Visit provider
Free accessFree model routes
Payment cardNo
AccountRequired
Sources11 first-party links
API endpointhttps://api.z.ai/api/paas/v4

Models mentioned

3
glm-4.7-flashglm-4.6v-flashglm-4.5-flash

Equivalent paid value

Not quantifiable

How this was valued: The published allowance or a defensible paid comparison is not precise enough to calculate.

The zero-price routes (glm-4.7-flash, glm-4.6v-flash, glm-4.5-flash) have documented per-request ceilings (200K context / 128K max output for GLM-4.7-Flash) and an exact-model paid reference exists (Z.AI's own accelerated GLM-4.7-FlashX at $0.07 input / $0.40 output per 1M, and OpenRouter's z-ai/glm-4.7-flash at $0.06/$0.40), but no finite period envelope survives either the documented or the modeled tier: neither docs.z.ai nor the operator's Chinese platform documentation publishes any numeric request-rate, token-rate, or concurrency ceiling for these routes — both state that per-model rates are visible only in the signed-in rate-limit console — so per-request ceilings alone leave requests per period unbounded. Modeling the missing request-rate quantity would have no first-party published example to cite (third-party reports of a 1-request concurrency cap are leads, not evidence, and a concurrency cap without a first-party generation-speed figure still yields no token envelope), which the modeled tier does not permit.

Limits and terms

Published Numeric Limits
No
Source Of Truth
Signed-in rate-limit dashboard.

What happens to your prompts?

PrivateReviewed

The API DPA describes real-time processing without storage and requires explicit agreement before content-based development or improvement.

Plan Scope
Z.AI developer API; the consumer-service terms differ materially.
Prompt Retention
The API DPA says prompt and generated content is processed in real time and not stored on Z.AI servers.
Response Retention
The API DPA gives generated API content the same no-storage treatment as input.
Ordinary Logging
Performance and usage metrics such as model version, inference, timing, diagnostics, and technical data may be collected and used to improve the service.
Model Training
API End User Content is not used to develop or improve services unless the API customer explicitly agrees.
Product Improvement
Technical usage metrics may be used for improvement; content use requires explicit agreement for API customers.
Human Or Operator Access
Content is processed to deliver the service and may be handled by approved subprocessors; no zero-operator-access promise is published.
Subprocessors And Routing
The DPA permits subcontractors and generally locates processing in Singapore; third-party model terms apply when selected.
Deletion Controls
Content is not stored under the API DPA; other customer data is deleted after termination unless law requires retention.
Caveat
Do not apply the consumer Z.AI policy—which permits broad content improvement use—to the developer API; the API Additional Terms and DPA control API content.

Eligibility

Account Required
Yes
Payment Method Required
not documented

Primary sources

11