Skip to main content

Runtime, approvals and safety

The Genie assistant runs on the OpenAI Codex CLI app-server, a separate program on your machine. Genie declares its own tools to the runtime, but it does not replace the runtime's built-in capabilities, and it does not override your runtime configuration. This page describes exactly what Genie sends, what it leaves to the runtime, and what that means for how you use it.

How the backend drives the runtime​

For each chat connection, the backend starts one runtime process (codex app-server, or the command set in GENIE_CODEX_COMMAND) and speaks JSON-RPC to it over standard input and output:

StepRequestWhat Genie sends
1initializeClient info (name: "genie", title: "Genie Chat") and the experimental-API capability
2thread/start (first message of a chat)The chosen model; the chat directory as working directory; base instructions; the mode's developer instructions; the mode's tool schemas as dynamic tools; approvalPolicy: null, sandbox: null, config: null
3turn/start (every message)Your text and pasted images, the model, the reasoning effort and a sandbox policy
—turn/interruptSent when you press Stop response
—thread/resume, thread/readSent when you reopen a chat, to resume the thread and replay its messages

The base instructions add one line to every chat: "ONLY work inside the current session working directory for session ... when creating or modifying files. Do not create files outside that session directory." This is an instruction to the model, not a control.

When the model calls a Genie tool, the runtime sends an item/tool/call request to the backend. The backend runs the tool in its own process, outside the runtime's sandbox, and returns the result. Everything the runtime emits, including every tool call with its arguments and result, is written to the chat's log, projects/<projectId>/sessions/<sessionId>/app-server.log.

If the runtime does not finish starting within 30 seconds, the chat reports "The assistant runtime did not finish initializing within 30 seconds."

Sandbox policy​

Every turn/start carries this sandbox policy:

{
"type": "workspaceWrite",
"networkAccess": true,
"writableRoots": ["<the chat's directory>"]
}

That is: workspace-write, network access on, and the chat's directory as the writable root. The policy applies to the runtime's own actions, such as shell commands and file edits. It turns network access on.

At thread start Genie sends no approval policy, no sandbox and no configuration (config: null), so the runtime fills them from its defaults and your configuration. In the logs of the reported sessions the runtime resolved the approval policy to "on-request", with approvals reviewer "auto_review".

Approvals: auto-accept, no approval UI​

The interface tells the backend to auto-accept approval requests (autoApprove: true). When the runtime asks to approve a command execution, a file change or a patch, the backend answers "accept" itself. Genie has no approval dialog.

In the reported sessions, no approval requests were raised.

Built-in tools are not disabled​

Genie does not disable the runtime's built-in tools. Besides the declared Genie tools, the agent can run shell commands, edit files in its writable area, use the runtime's web search and call tools from MCP servers in your runtime configuration. Shell commands, file edits and MCP calls appear in the chat as tool cards; web search does not, and is recorded only in the log. None of these produce Genie citation tokens.

The declared tools are not an enforced boundary

In the logged test sessions, the agent used the runtime's built-in web search in 2 of the 4 sessions (BCL11A and the FTO/IRX3 obesity locus), and in the FTO/IRX3 session also a shell curl command to the UCSC REST API. No approval request was raised, so disabling the backend's automatic acceptance of approval requests alone would not have stopped these actions. In the BCL11A session, a literature reference the agent gave had the wrong first author. The MYC and ALB demonstration sessions used only declared tools.

Do not read the Tool reference as a list of everything the agent can do. It is the list of what Genie declares and executes, and it is the only part that issues citations.

Your runtime configuration is loaded​

Because Genie sends config: null, the runtime uses your own user-level configuration: sign-in and model provider, user-level instruction files, and MCP servers. These can change how the agent behaves in Genie, for example how it cites or which extra tools it has.

In every logged session, the runtime loaded a user-level instruction file from the developer's configuration and started MCP servers from that configuration; no MCP tool was called. The instruction file's contents are not reported here, and the sessions should be rerun with it disabled.

Model list and reasoning effort​

  • Models. The model picker lists what the local CLI reports (model/list). The backend asks a short-lived runtime process and caches the list for 5 minutes. If the list cannot be fetched and nothing is cached, Genie shows a fallback list: GPT-5.5 (default), GPT-5.4 and GPT-5.4-Mini.
  • Reasoning effort. Low, Medium, High or Extra High (sent as low, medium, high, xhigh). Your choice is stored per chat and sent with every turn. The runtime's thread-level default is "medium", but each turn overrides it. Asset-import chats always send "medium".
Reported sessionsModel, effortCLI version
The earlier sessions (screenshots only)composer label "GPT-5.6-Sol · Low"not recorded
The four test sessionsgpt-5.6-sol, low0.150.1
The two demonstration sessions (MYC, ALB)gpt-6.1-sol, low0.159.1

Between the earlier sessions and the demonstration sessions, every tool value visible in both matched at the displayed precision. That shows the tools returned the same numbers for the same inputs (tool reproducibility); it is not evidence that the models agree.

Reading the logs​

The chat's app-server.log is the full record of a session. These logs were used to check every tool-computed number in the logged sessions and to find actions outside the declared tools. To review a session:

  • find the Genie tool calls by name (for example public_data_search or genome_call_peaks) and read their arguments and results;
  • look for runtime items of type webSearch, commandExecution and mcpToolCall, which are actions outside the declared tools;
  • look for approval requests (requestApproval), which the backend accepts automatically;
  • compare numbers in the agent's answer with the tool results.

Recommendations​

  • Check the citations. Open each chip and compare the record with the sentence. A chip shows which record a claim points to, not that the sentence is correct. See Citations and evidence.
  • Read the tool cards and the logs. Check search relaxations, absolute peak cutoffs, bin sizes and any error. Look in the log for web searches and shell commands you did not expect.
  • Run in a dedicated environment. The agent can run shell commands with network access and auto-accepted approvals. Use a machine or user account set aside for this work, without credentials or data you would not want a shell command to reach.
  • Keep your runtime configuration minimal. Remove instruction files and MCP servers you do not need for Genie, or use a separate runtime configuration for it.
  • Keep the backend local. The local backend has no sign-in of its own and listens on 127.0.0.1 by default. Do not expose it on a network.
  • Treat interpretations as the model's own. In logged test sessions, the agent gave correct numbers followed by unsupported inferences. See Limitations and known issues.

Suggested design changes​

The following changes have been suggested. None of them had been implemented when the logged sessions were run.

  • Do not present one window's correlation as agreement: report how r depends on bin size and method (for example r at several bin sizes and a rank correlation), and for a question about agreement offer a design that can answer it.
  • Keep each track's assay and ENCODE output type with the loaded track and show them next to r.
  • Merge peak intervals into regions and set a minimum width.
  • Cite viewport statistics to a record that holds the numbers.
  • Require the agent to report search relaxations, and show them in the interface.
  • Gate runtime tools outside the declared set (built-in web search, shell) through the approval policy, sandbox and runtime configuration that Genie sends.