For AI agents: markdown of this page — /docs-content-en/changelog/2026-07-06.md documentation index — /llms.txt

API changes: July 6, 2026

← Changelog · July 2026

NEW-0706-1: pricing.perCall and pricing.perMinute fields in the model catalog

GET /v1/models and GET /v1/models/{model} responses gain optional fields in the pricing object: perCall — the cost of a single request in Vibe credits, perMinute — the cost of one minute of audio in Vibe credits (for speech-to-text models). The fields appear only for models whose corresponding base price is above zero; for all other models the pricing object is unchanged — existing requests keep working as before.

NEW-0706-2: New 402 error code ai_quota_exhausted on AI endpoints

When the account monthly AI quota control is active, POST /v1/chat/completions, POST /v1/embeddings and POST /v1/audio/transcriptions may return 402 with { success: false, error: { code: "ai_quota_exhausted", type: "insufficient_quota", reason, resetAt? } }. The reason field distinguishes three cases: breaker — the hourly over-quota spending limiter fired, wallet_empty — the quota is exhausted and the account balance has no funds, wallet_off — over-quota usage is not available for this account. resetAt is when requests will pass again (may be absent for a permanently disabled account). While the account quota is not exhausted, endpoint behavior is unchanged.

FIX-0706-3: Over-quota AI usage is charged at the model's base catalog price

Before

With the account monthly AI quota control active, over-quota requests to POST /v1/chat/completions, POST /v1/embeddings and POST /v1/audio/transcriptions were charged to the account balance at internal quota-program rates — with off-peak discounts applied; the effective price was not visible in the model catalog.

After

Over-quota usage is charged at the model's base price from the public catalog — the same one returned in the pricing field of GET /v1/models, including the new perCall and perMinute for non-token models. Off-peak discounts apply only to quota consumption, not to the money balance. Usage within the quota is still not charged to the account balance.

Impact on integrators

No changes required. The cost of over-quota usage can now be computed upfront from the model's catalog price.

NEW-0706-4: The bitrix/embeddings model is available in the API

The POST /v1/embeddings endpoint is now served by the bitrix/embeddings model — turning text into vector representations for semantic search, clustering, and retrieval (RAG). The model is free and billed on input tokens only. For the list of embedding-capable models see GET /v1/models.

FIX-0706-5: Deploy reports an honest error when the new build did not bind the port

Deploy via POST /v1/infra/servers/:id/deploy now verifies the port is held by the NEW service. If a previous process keeps listening and the new build crash-loops with EADDRINUSE, the deploy fails honestly instead of falsely succeeding; a port held by a leftover process of the same app is freed automatically where provable.

Before

The old version kept answering 200, the deploy reported success, and the new build never came up — with no error and no hint.

After

The healthcheck step returns an error naming EADDRINUSE and the port, and the stop_existing step frees the port from a leftover process of the app (or warns and proceeds when it cannot).

NEW-0706-6: filter: $nin (NOT IN) operator to exclude a set of values

Before

There was no way to select records whose field is NOT in a set of values: the $in (IN) operator existed, but its inverse did not. The Bitrix24-native field-name prefixes @ (IN) and !@ (NOT IN) ({ "!@categoryId": [1, 3] }) were not translated — such a deal filter returned 400 UNKNOWN_FILTER_FIELD.

After

Added the $nin operator: { "filter": { "categoryId": { "$nin": [1, 3] } } } returns records of every value except those listed (NOT IN). Symmetric to $in. The Bitrix24-native @ / !@ field-name prefixes are still unsupported, but now return a clear 400 INVALID_FILTER_FIELD hinting to switch to $in / $nin, instead of a confusing error or a silently-ignored (thus full-set) filter.

BC-0706-7: dedicated error code for an oversized exec command

Old format supported until: 06.01.2027

Before

A command longer than 10000 characters at POST /v1/infra/servers/:id/exec was rejected with the generic VALIDATION_ERROR code, with no cause and no way out.

After

Such a request returns 400 with the dedicated COMMAND_TOO_LONG code and a structured hint: ship large payloads and scripts via POST /v1/infra/servers/:id/upload, then run them with bash /path/script.sh. All other schema violations still return VALIDATION_ERROR.

What integrators should do

If your client handles VALIDATION_ERROR of this endpoint as the catch-all validation case — add handling for the COMMAND_TOO_LONG code (or treat any 400 uniformly).

NEW-0706-8: hint in the exec timeout error

The EXEC_TIMEOUT error of POST /v1/infra/servers/:id/exec now carries a structured hint field (reason / recovery / recoveryAction): why the process was stopped (at timeout the whole process group is terminated forcibly, with no grace period) and what to do — run long operations as a detached background job and monitor it via GET /v1/infra/servers/:id/logs, raise timeout up to 600 seconds, or use ?stream=true. The field is additive: the previous code / message shape is unchanged, and the hint arrives both in JSON mode and in the SSE error event.