Skip to main content

Architecture

Genie pairs a genome browser, rendered by eg3 (the third-generation WashU Epigenome Browser engine), with an LLM agent that drives the browser through declared tools. Genie's own code is the workspace around the browser, the backend that runs the tools, and the citation layer that ties answers to stored tool results.

Components​

ComponentCodeRole
Rendererapps/standaloneVite + React 19 + React Router 7 single-page app. Project list, Genome mode, Analysis mode, shared read-only view. Runs in a browser tab or in the Electron window.
eg3 enginevendor/eg3 (@eg3/adapter, @eg3/tracks)Renders the genome browser panel in project and shared views. See Browser engine.
Chatpackages/chat-core, packages/chat-ui, packages/codex-runtimeChat state, the ChatView component with Markdown and citation chips, and the WebSocket transport to the backend.
Backendapps/backend, packages/backend-coreNode HTTP server and WebSocket endpoint /ws. Stores projects, sessions, assets and browser state; runs Genie tools; bridges to the Codex CLI. See Backend reference.
Codex app-serverExternal codex CLIAgent runtime. The backend starts codex app-server for each WebSocket connection and talks JSON-RPC over standard input and output. The model is chosen at run time from the models the local CLI lists.
Session datagenie/session-data/ (or Electron's user-data folder)JSON files, per-project folders, result files, per-chat logs and citation stores.
Cloud project APIinfra/aws-cdk/lambda/api, or the genie-local-cloud simulatorProject REST API for cloud mode and sharing. No assistant. See Desktop and cloud.
ENCODE portalhttps://www.encodeproject.org/Searched by the public_data_search tool; its file URLs are loaded as tracks.
Gene annotation APIhttps://lambda.epigenomegateway.org/v3Gene tracks and the "Gene symbol" search box in the UI. The agent has no gene lookup tool.
Public data hubsvizhub.wustl.edu hub manifestsListed in the Public Data Hubs sheet.

The agent-to-browser control loop​

The agent never calls into the renderer directly. Genie tools run on the backend and change the project's browser state file; the frontend notices and reloads.

  1. The frontend sends the user's message over the WebSocket. The backend starts a turn on the Codex thread.
  2. When the model calls a Genie tool, the app-server sends an item/tool/call request. CodexSession.handleServerRequest runs the handler on the backend.
  3. Mutating genome tools (genome_navigate, genome_add_track, genome_call_peaks, genome_update_builtin_track, and the disabled custom-track tools) write projects/<projectId>/genome/browser-state.json.
  4. The backend answers the app-server with the tool output and sends tool_call and tool_result events to the frontend.
  5. The workspace watches tool-call events. When a call to one of the mutating genome tools finishes, it bumps a refresh key and the genome browser panel refetches GET /projects/:projectId/genome. Read-only tools (statistics, search, listing) do not trigger a refetch.

User edits go the other way: the browser panel saves track and view changes with PATCH /projects/:projectId/genome, and the view region is saved after a 150 ms debounce. Tools read the same file, so the agent always works from the last saved state.

Data flow for a computation tool​

Computation tools do not look at the rendered image. They read the underlying data files on the backend.

Take genome_signal_stats as an example:

  1. The handler loads browser-state.json, finds the requested track ids and the current view region, and resolves the scope (viewport, chromosome, all, or explicit regions) to a list of loci.
  2. If the loci add up to more than 5,000,000 bases, the tool returns a "scope too large for synchronous analysis" result instead of fetching.
  3. genome/regionData.js opens the bigWig file with @gmod/bbi and a generic-filehandle RemoteFile, which reads over HTTP range requests, and reads the intervals that overlap each locus. bedGraph tracks are read as text or through tabix. Chromosome names are mapped to Ensembl style (no chr prefix) when the track or genome needs it.
  4. The handler summarizes the intervals (minimum, maximum, length-weighted mean, sum, covered fraction).
  5. The result is decorated with a citation (see below) and returned to the model as JSON.

genome_call_peaks and genome_quantify_at_features also write result files to genome/results/ (BED and TSV). Peak calling adds the BED file as a new track in the browser state. See Measurements for the method of each tool.

Citation pipeline​

Citations are issued by the tools, not written by the model.

  1. A tool that produces evidence creates a citation record with an id cite_... and returns, inside its JSON result, a citationToken ([[cite:<id>]]), citationInstructions and a citations array.
  2. The backend adds the records to the chat's citations.json (schema version 1) in the session folder.
  3. The mode's developer instructions tell the model to copy the token right after the claim it supports, never to invent ids, and to mark uncited claims.
  4. ChatView finds tokens with the pattern \[\[cite:([a-zA-Z0-9_-]+)\]\]. It builds a registry from message metadata and from tool-call outputs. Tokens whose id is in the registry become numbered chips (numbering restarts in each message); unknown ids stay as raw text.
  5. Selecting a chip opens the evidence panel, which shows the stored record: source type, title, location, token, genome region and an excerpt (computation results are clamped to 1,800 characters), with an "Open source" link when the record has a URL.
  6. When a chat is reopened, the backend replays the thread through thread/read and reattaches the stored citations to the replayed messages.
Source typeIssued by
genome-regiongenome_get_current_region, genome_navigate, genome_describe_viewport (region only)
genome-computationgenome_signal_stats, genome_call_peaks, genome_quantify_at_features, genome_correlate_tracks
public-data-search, public-data-filepublic_data_search (one record for the search, one per ENCODE file)
asset-catalog, asset, remote-sourceProject asset tools
studio-outputproject_read_analysis_plan

A chip shows which record a claim points to. It does not show that the sentence is correct, and uncited text is not blocked. See Citations and evidence.

Two rendering engines​

EngineUsed inPackages
eg3 (vendored)Project workspace (Genome mode) and read-only shared views (/shared/:shareId)@eg3/adapter, @eg3/tracks in vendor/eg3
Native packagesPublic landing browser in static builds (/ and /browser/:genomeId, hg38 and hg19), developer pages under /dev/genome-track*, apps/genome-track-demo, and the IGVF bundle built by pnpm bundle:igvf-genome@genie/genome-browser, @genie/genome-track-container, @genie/genome-tracks-standard

The workspace also uses the native packages for types and region models (TrackModel, DisplayedRegionModel, NavigationContext, ChromosomeInterval) for locus math, persistence and gene search. See Browser engine.

Frontend state​

ConcernLibraryNotes
Server dataTanStack QueryProjects, browser state (GET/PATCH /projects/:id/genome), assets, Studio outputs, models
Global UI statezustandStores under apps/standalone/src/app/stores/
Notifications and errorssonnerToasts such as "Genome browser image exported"
Chat state@genie/chat-coreChatStore and ChatSession, fed by WebSocketChatTransport
Layout splitlocalStorage key genie.workspace-uiPer-browser preference

Desktop, local web and cloud​

The same renderer serves three deployments:

  • Desktop: Electron starts the backend in-process on 127.0.0.1 and loads the renderer from it. Assistant on.
  • Local web: Vite dev server plus a separate backend on :8787. Assistant on only when enabled.
  • Cloud: static files on S3 and CloudFront, Lambda project API, Cognito sign-in. No assistant.

See Desktop and cloud and Configuration.