# API changes: July 6, 2026

[← Changelog](/docs/changelog) · [July 2026](/docs/changelog/2026-07)

### NEW-0706-1: pricing.perCall and pricing.perMinute fields in the model catalog

[GET /v1/models](/docs/ai/models/list) and [GET /v1/models/{model}](/docs/ai/models/get) responses gain optional fields in the `pricing` object: `perCall` — the cost of a single request in Vibe credits, `perMinute` — the cost of one minute of audio in Vibe credits (for speech-to-text models). The fields appear only for models whose corresponding base price is above zero; for all other models the `pricing` object is unchanged — existing requests keep working as before.

### NEW-0706-2: New 402 error code ai_quota_exhausted on AI endpoints

When the account monthly AI quota control is active, [POST /v1/chat/completions](/docs/ai/chat/completions), [POST /v1/embeddings](/docs/ai/embeddings) and [POST /v1/audio/transcriptions](/docs/ai/audio/transcriptions) may return 402 with `{ success: false, error: { code: "ai_quota_exhausted", type: "insufficient_quota", reason, resetAt? } }`. The `reason` field distinguishes three cases: `breaker` — the hourly over-quota spending limiter fired, `wallet_empty` — the quota is exhausted and the account balance has no funds, `wallet_off` — over-quota usage is not available for this account. `resetAt` is when requests will pass again (may be absent for a permanently disabled account). While the account quota is not exhausted, endpoint behavior is unchanged.

### FIX-0706-3: Over-quota AI usage is charged at the model's base catalog price

**Before**

With the account monthly AI quota control active, over-quota requests to [POST /v1/chat/completions](/docs/ai/chat/completions), [POST /v1/embeddings](/docs/ai/embeddings) and [POST /v1/audio/transcriptions](/docs/ai/audio/transcriptions) were charged to the account balance at internal quota-program rates — with off-peak discounts applied; the effective price was not visible in the model catalog.

**After**

Over-quota usage is charged at the model's base price from the public catalog — the same one returned in the `pricing` field of [GET /v1/models](/docs/ai/models/list), including the new `perCall` and `perMinute` for non-token models. Off-peak discounts apply only to quota consumption, not to the money balance. Usage within the quota is still not charged to the account balance.

**Impact on integrators**

No changes required. The cost of over-quota usage can now be computed upfront from the model's catalog price.

### NEW-0706-4: The bitrix/embeddings model is available in the API

The [POST /v1/embeddings](/docs/ai/embeddings) endpoint is now served by the `bitrix/embeddings` model — turning text into vector representations for semantic search, clustering, and retrieval (RAG). The model is free and billed on input tokens only. For the list of embedding-capable models see [GET /v1/models](/docs/ai/models/list).

### FIX-0706-5: Deploy reports an honest error when the new build did not bind the port

Deploy via [POST /v1/infra/servers/:id/deploy](/docs/infra/deploy) now verifies the port is held by the NEW service. If a previous process keeps listening and the new build crash-loops with EADDRINUSE, the deploy fails honestly instead of falsely succeeding; a port held by a leftover process of the same app is freed automatically where provable.

**Before**

The old version kept answering `200`, the deploy reported success, and the new build never came up — with no error and no hint.

**After**

The `healthcheck` step returns an error naming EADDRINUSE and the port, and the `stop_existing` step frees the port from a leftover process of the app (or warns and proceeds when it cannot).

### NEW-0706-6: filter: $nin (NOT IN) operator to exclude a set of values

**Before**

There was no way to select records whose field is NOT in a set of values: the `$in` (IN) operator existed, but its inverse did not. The Bitrix24-native field-name prefixes `@` (IN) and `!@` (NOT IN) (`{ "!@categoryId": [1, 3] }`) were not translated — such a deal filter returned `400 UNKNOWN_FILTER_FIELD`.

**After**

Added the `$nin` operator: `{ "filter": { "categoryId": { "$nin": [1, 3] } } }` returns records of every value except those listed (NOT IN). Symmetric to `$in`. The Bitrix24-native `@` / `!@` field-name prefixes are still unsupported, but now return a clear `400 INVALID_FILTER_FIELD` hinting to switch to `$in` / `$nin`, instead of a confusing error or a silently-ignored (thus full-set) filter.

### BC-0706-7: dedicated error code for an oversized exec command

> Old format supported until: 06.01.2027

**Before**

A command longer than 10000 characters at [POST /v1/infra/servers/:id/exec](/docs/infra/deploy/exec) was rejected with the generic `VALIDATION_ERROR` code, with no cause and no way out.

**After**

Such a request returns 400 with the dedicated `COMMAND_TOO_LONG` code and a structured `hint`: ship large payloads and scripts via [POST /v1/infra/servers/:id/upload](/docs/infra/deploy/upload), then run them with `bash /path/script.sh`. All other schema violations still return `VALIDATION_ERROR`.

**What integrators should do**

If your client handles `VALIDATION_ERROR` of this endpoint as the catch-all validation case — add handling for the `COMMAND_TOO_LONG` code (or treat any 400 uniformly).

### NEW-0706-8: hint in the exec timeout error

The `EXEC_TIMEOUT` error of [POST /v1/infra/servers/:id/exec](/docs/infra/deploy/exec) now carries a structured `hint` field (`reason` / `recovery` / `recoveryAction`): why the process was stopped (at `timeout` the whole process group is terminated forcibly, with no grace period) and what to do — run long operations as a detached background job and monitor it via [GET /v1/infra/servers/:id/logs](/docs/infra/deploy/logs), raise `timeout` up to 600 seconds, or use `?stream=true`. The field is additive: the previous `code` / `message` shape is unchanged, and the hint arrives both in JSON mode and in the SSE `error` event.
