Z.AI Model API
3 model routes currently listed at zero price.
https://api.z.ai/api/paas/v4Models mentioned
glm-4.7-flashglm-4.6v-flashglm-4.5-flashEquivalent paid value
How this was valued: The published allowance or a defensible paid comparison is not precise enough to calculate.
The zero-price routes (glm-4.7-flash, glm-4.6v-flash, glm-4.5-flash) have documented per-request ceilings (200K context / 128K max output for GLM-4.7-Flash) and an exact-model paid reference exists (Z.AI's own accelerated GLM-4.7-FlashX at $0.07 input / $0.40 output per 1M, and OpenRouter's z-ai/glm-4.7-flash at $0.06/$0.40), but no finite period envelope survives either the documented or the modeled tier: neither docs.z.ai nor the operator's Chinese platform documentation publishes any numeric request-rate, token-rate, or concurrency ceiling for these routes — both state that per-model rates are visible only in the signed-in rate-limit console — so per-request ceilings alone leave requests per period unbounded. Modeling the missing request-rate quantity would have no first-party published example to cite (third-party reports of a 1-request concurrency cap are leads, not evidence, and a concurrency cap without a first-party generation-speed figure still yields no token envelope), which the modeled tier does not permit.
Limits and terms
- Published Numeric Limits
- No
- Source Of Truth
- Signed-in rate-limit dashboard.
What happens to your prompts?
The API DPA describes real-time processing without storage and requires explicit agreement before content-based development or improvement.
- Plan Scope
- Z.AI developer API; the consumer-service terms differ materially.
- Prompt Retention
- The API DPA says prompt and generated content is processed in real time and not stored on Z.AI servers.
- Response Retention
- The API DPA gives generated API content the same no-storage treatment as input.
- Ordinary Logging
- Performance and usage metrics such as model version, inference, timing, diagnostics, and technical data may be collected and used to improve the service.
- Model Training
- API End User Content is not used to develop or improve services unless the API customer explicitly agrees.
- Product Improvement
- Technical usage metrics may be used for improvement; content use requires explicit agreement for API customers.
- Human Or Operator Access
- Content is processed to deliver the service and may be handled by approved subprocessors; no zero-operator-access promise is published.
- Subprocessors And Routing
- The DPA permits subcontractors and generally locates processing in Singapore; third-party model terms apply when selected.
- Deletion Controls
- Content is not stored under the API DPA; other customer data is deleted after termination unless law requires retention.
- Caveat
- Do not apply the consumer Z.AI policy—which permits broad content improvement use—to the developer API; the API Additional Terms and DPA control API content.
Governing documents
- Terms Of Service
- https://docs.z.ai/legal-agreement/terms-of-use
- Privacy Policy
- https://docs.z.ai/legal-agreement/privacy-policy
Eligibility
- Account Required
- Yes
- Payment Method Required
- not documented