> ## Documentation Index
> Fetch the complete documentation index at: https://docs.onecortex.io/llms.txt
> Use this file to discover all available pages before exploring further.

# The invoke API

> Call a deployed agent over HTTPS: the request, the JSON response, the event stream, the headers, and every error it can answer.

Every deployed agent answers one HTTPS endpoint. You send a prompt, and you get back either one JSON body or a stream of server sent events carrying the agent's text, tool calls, tool results and steps as they happen.

## The endpoint

Every agent has one endpoint:

```text theme={null}
POST https://api.onecortex.io/v1/agents/{agentId}/invoke
```

The agent's page in the dashboard shows the full URL and the agent's ID, which starts with `agt_`, each with a copy button.

## Authentication

Every call carries an API key as a bearer token:

```http theme={null}
Authorization: Bearer $ONECORTEX_API_KEY
```

Create a key under **API keys** in the dashboard. It is shown once, when you create it. See [API keys](/call/api-keys).

A key that is missing, unknown, revoked, expired, or scoped to a different agent is refused with `401` and the same message, `A valid Onecortex API key is required.`, except the last case, which answers `404` so the key never learns that the other agent exists.

## The request

A JSON body, at most 1 MB.

<ParamField body="prompt" type="string" required>
  The input to your agent. A request without it is refused with `400` before your agent is called: `Request must include a 'prompt' field.`
</ParamField>

<ParamField body="sessionId" type="string">
  Your own key for a conversation, 1 to 256 characters. Send the same value again to reach the same running instance of your agent. Leave it out and Onecortex generates one and returns it. See [Sessions](/build/sessions).
</ParamField>

<ParamField body="stream" type="boolean">
  `true` for server sent events, `false` for one JSON body. Left out, the `Accept` header decides: `text/event-stream` streams, anything else (including `*/*`) returns JSON.
</ParamField>

<ParamField body="messages" type="array">
  The conversation as you hold it, for an agent that reads history. Each item is `{ "role": "...", "content": "..." }`.
</ParamField>

<ParamField body="metadata" type="object">
  Anything you want your agent to see, at most 4 KB as JSON. Larger is refused with `400`: `The 'metadata' field can be at most 4 KB.`
</ParamField>

Every other field you send reaches your agent unchanged, as `params`, with `metadata` among them. Onecortex does not decide which of your fields mean something. Each [framework page](/frameworks/langgraph) says where `params` arrive in your code.

```json Request body theme={null}
{
  "prompt": "hello",
  "sessionId": "user-42",
  "stream": false,
  "model": "gpt-5",
  "metadata": { "tenant": "acme" }
}
```

Here, your agent receives `params` of `{"model": "gpt-5", "metadata": {"tenant": "acme"}}`.

## The JSON response

With `stream` false, the response is one body once the run is over.

<ResponseField name="status" type="string">
  `completed` or `failed`.
</ResponseField>

<ResponseField name="result" type="string">
  The whole reply. On a failed run, whatever text arrived before it failed.
</ResponseField>

<ResponseField name="events" type="array">
  The tool calls, tool results and steps, in order, as [events](/build/streaming). Text deltas are left out, because `result` holds them whole.
</ResponseField>

<ResponseField name="error" type="object">
  Present only when `status` is `failed`: `{ "code": "...", "message": "..." }`.
</ResponseField>

<ResponseField name="sessionId" type="string">
  Your `sessionId`, echoed unchanged, or the one generated for you.
</ResponseField>

<ResponseField name="agentId" type="string">
  The agent that answered.
</ResponseField>

<ResponseField name="versionNumber" type="number">
  The version that answered, as shown on the agent's **Versions** tab.
</ResponseField>

<ResponseField name="requestId" type="string">
  This request's ID, starting `req_`. Quote it to support.
</ResponseField>

<ResponseField name="durationMs" type="number">
  How long your agent took, in milliseconds.
</ResponseField>

For the `echo` agent from the [quickstart](/quickstart):

```bash Terminal theme={null}
curl https://api.onecortex.io/v1/agents/agt_.../invoke \
  -H "Authorization: Bearer $ONECORTEX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"hello"}'
```

```json Response theme={null}
{
  "status": "completed",
  "result": "You said: hello",
  "events": [
    { "v": 1, "type": "tool_call_start", "id": "call_1", "name": "word_count" },
    { "v": 1, "type": "tool_call_args", "id": "call_1", "delta": "{\"text\": \"hello\"}" },
    { "v": 1, "type": "tool_call_end", "id": "call_1" },
    { "v": 1, "type": "tool_result", "id": "call_1", "output": "1" }
  ],
  "sessionId": "01K...",
  "agentId": "agt_...",
  "versionNumber": 1,
  "requestId": "req_01K...",
  "durationMs": 212
}
```

### When your agent raises

If your agent's own code raises, the run fails, and the response is `502` with the same body: `status` is `failed`, `result` holds any text sent before the failure, and `error` names the exception. Your traceback is not in the response. It is in your agent's [logs](/observe/logs).

```json Response (502) theme={null}
{
  "status": "failed",
  "result": "",
  "error": { "code": "agent_error", "message": "ValueError: the agent was asked to raise" },
  "events": [],
  "sessionId": "01K...",
  "agentId": "agt_...",
  "versionNumber": 1,
  "requestId": "req_01K...",
  "durationMs": 31
}
```

## The event stream

With `stream` true, the response is `text/event-stream`. Each event is one frame: an `event:` line naming its type, a `data:` line holding the event as JSON on a single line, and a blank line.

```text Stream theme={null}
event: text
data: {"v":1,"type":"text","delta":"You said: hello"}

event: done
data: {"v":1,"type":"done","result":"You said: hello","sessionId":"01K...","versionNumber":1,"durationMs":212}
```

* Exactly one terminal event ends every stream: `done` or `error`. Nothing follows it.
* `done` carries the whole reply in `result`, plus `sessionId`, `versionNumber` and `durationMs`.
* While your agent is working and has nothing to send, Onecortex writes a comment line, `: ping`, every 15 seconds, so no proxy closes an idle connection. Server sent event clients ignore comments.

Every event type and its fields are on [Streaming and events](/build/streaming).

<Warning>
  A stream that has started has already sent `200 OK`, and cannot change it. If the run fails after that, the failure arrives in the stream as `event: error`. Handle it: a client that reads only `text` and `done` shows a failed run as an empty success.
</Warning>

```text Stream that fails theme={null}
event: error
data: {"v":1,"type":"error","code":"agent_error","message":"ValueError: the agent was asked to raise"}
```

## Response headers

| Header | On | Holds |
| - | - | - |
| `X-Request-Id` | Every response | The request ID, `req_...`, also in every error body |
| `X-Onecortex-Session-Id` | Every response that reached your agent | The session ID, so a caller that sent none can continue the conversation |
| `X-RateLimit-Remaining` | Every response that reached your agent | Requests left in your organization's burst allowance |
| `Retry-After` | `429` | Whole seconds to wait before trying again |

## Errors

An error before the run starts is an HTTP status with this body:

```json theme={null}
{
  "error": {
    "code": "not_found",
    "message": "That agent does not exist.",
    "requestId": "req_01K..."
  }
}
```

| Status | `code` | Message | Retry |
| - | - | - | - |
| 400 | `invalid_request` | `Request must include a 'prompt' field.` | No, fix the request |
| 400 | `invalid_request` | `The request body must be valid JSON.` | No, fix the request |
| 400 | `validation_failed` | Names the field, for example the 4 KB `metadata` limit | No, fix the request |
| 401 | `unauthenticated` | `A valid Onecortex API key is required.` | No, check the key |
| 403 | `organization_suspended` | `This organization is suspended. Contact Onecortex support.` | No |
| 404 | `not_found` | `That agent does not exist.` | No, check the ID and the key's scope |
| 409 | `agent_not_ready` | `This agent has not finished its first deployment yet.` | Yes, once the first build succeeds |
| 409 | `agent_not_deployed` | `This agent has no deployed version. Its first deployment did not succeed.` | No, fix the build and deploy |
| 410 | `agent_deleted` | `This agent has been deleted.` | No |
| 413 | `payload_too_large` | `The request body can be at most 1 MB.` | No |
| 429 | `rate_limited` | `Too many requests. Try again in <n> seconds.` | Yes, after `Retry-After` |
| 502 | `agent_error` | Your agent's exception, as `<Type>: <message>` | Only if your code would succeed next time |
| 502 | `upstream_error` | `The agent could not be reached. Try again.` | Yes |
| 503 | `agent_unavailable` | `This agent is not available right now. Try again.` | Yes |
| 503 | `service_unavailable` | `Onecortex is temporarily unavailable. Try again.` | Yes |
| 503 | `capacity_exceeded` | `Onecortex is at capacity in this region.` | Yes, later |

An agent still serves while a new version builds, and while a new version's build has failed: a deploy never takes a working agent offline. Every error, with its cause and fix, is on [Errors](/production/errors).

## Limits

60 requests a minute per organization with a burst of 20, 30 a minute per agent, and 10 streams open at once per organization. A request body is at most 1 MB and `metadata` at most 4 KB. All limits are on [Limits](/production/limits).

## Examples in your language

<Columns cols={3}>
  <Card title="Python" icon="file-code" href="/call/python">
    `httpx`, streamed.
  </Card>

  <Card title="TypeScript" icon="file-code" href="/call/typescript">
    `fetch`, with the stream reader.
  </Card>

  <Card title="curl" icon="terminal" href="/call/curl">
    `curl -N`, from a terminal.
  </Card>
</Columns>
