Limitations and known issues
This page collects the limits of Genie and of the evidence about it, plus known app defects. Read it before relying on an agent answer for a research conclusion.
Limits of the evidence
The evidence from logged sessions is small and was gathered by the developers:
- Few sessions. Two logged demonstration sessions (at the MYC and ALB loci), four logged test sessions, and two earlier sessions known only from screenshots. A third earlier session, at rs12740374, is excluded because its measured interval misses the variant. All are on hg38, and mostly H3K27ac in K562 or HepG2 cells.
- No benchmark, rate or study. There is no task benchmark, no citation-validity or hallucination rate, no user study, no timing measurement and no tool-free baseline. Genie makes no speed or accuracy claims.
- Run by the development team. All sessions used prepared prompts written by the team, and every navigation prompt gave coordinates.
- Automated reviews. The reviews of the test sessions were automated LLM fact-checks, not domain-expert reviews.
- One portal, one runtime. Genie searches only ENCODE, and runs on one agent runtime (the OpenAI Codex CLI app-server) and one model provider.
- Hidden configuration. Every logged session also loaded a developer-local instruction file whose contents are not reported and could have shaped the agent's behaviour, for example its citation compliance. A rerun with it disabled is needed. See Runtime and safety.
What citations do and do not guarantee
A numbered chip in the chat means that the agent copied a citation token that a Genie tool issued in that chat, and that the token resolves to a stored record. See Citations and evidence.
A chip does not guarantee that:
- The sentence is supported. A resolving citation is not a supporting citation. A chip shows which record a claim points to, not that the sentence is correct. In a logged test session, a correct correlation result was cited for an inference it did not support.
- Every claim is cited. Nothing blocks uncited text. The agent is instructed to mark uncited claims, but the app does not enforce it.
- The record holds the number. Viewport statistics (means, maxima, coverage) are cited to a region record that holds only the locus. This happened in both demonstration answers. The numbers are in the session log, not in the evidence panel.
- Outside knowledge is checked. Anything the agent learns outside Genie's tools, for example through the runtime's web search, is outside the citation system. In one test session a literature citation named the wrong first author.
Runtime boundary
Genie's declared tools are not an enforced boundary:
- Genie does not disable the runtime's built-in web search or shell. In the logged test sessions, the agent used web search in 2 of 4 sessions and a shell
curlcommand in 1. - Each turn runs in a workspace-write sandbox with network access on.
- The backend is configured to auto-accept the runtime's approval requests, but no approval request was raised for those actions. Turning off auto-accept alone would not have stopped them.
- The runtime can also start MCP servers from the user's own runtime configuration. In the reported sessions they started but were never called.
See Runtime and safety.
Tool limits
5 Mb scope limit
The four computation tools (genome_signal_stats, genome_call_peaks, genome_quantify_at_features, genome_correlate_tracks) refuse scopes larger than 5,000,000 bp with the message "scope too large for synchronous analysis ... narrow the region or use viewport scope". On hg38 every chromosome except chrM is larger than 5 Mb, so chromosome-wide and genome-wide scopes fail. genome_describe_viewport has no such check, but above 5 Mb its per-track signal fetches fail. There are no background jobs or cancellation. (The 5,000,000 bp limit is still in the current code.)
The peak caller is a threshold caller
genome_call_peaks keeps fetched intervals scoring at least the mean plus 2 standard deviations of all fetched scores (by default), merges only overlapping or touching intervals, and keeps intervals as short as 1 bp. Its own note says: "Simple z-score threshold caller over fetched intervals; not a model-based peak caller." It uses no control track, p-value, false discovery rate or blacklist. Consequences:
- Peak counts are fragment counts. One enriched region is usually reported as several intervals. In the ALB demonstration, 32 HepG2 intervals formed about 4 regions and 4 K562 intervals one region.
- Cutoffs are relative to the window. The cutoff comes from the signal in the current scope, so "no call" at a position does not mean there is no signal there.
- The output field
thresholdholds the absolute cutoff, not a z value.
Correlation has no normalization check
genome_correlate_tracks computes a binned Pearson r over one scope. It does not check that the two tracks have the same assay or output type, reports no significance value, and gives a value whose size can depend strongly on bin size and on a few bins. One window's r is not a measure of agreement between assays or of tissue specificity. See Measurements for the GATA1 example.
Search is ENCODE-only and can relax its filters
public_data_search queries only the ENCODE portal. Its first file query asks for "fold change over control" signal; if nothing is found it relaxes the query up to three times (drop the output type, then match the biosample by organ, then by cell type). Relaxations are returned in the result, but the agent may not mention them, as happened in a logged test session. Results come from the live portal and can change over time. See Public data search.
The agent cannot look up genes by symbol
The navigation tool resolves only coordinates (chr:start-end); a gene symbol does not move the viewport to the gene, even though the tool schema mentions gene symbols (the symbol is stored without coordinates and the browser falls back to the genome's default region). No tool lists the genes in view, and the agent cannot see the rendered browser. Give the agent coordinates. (The Gene symbol search box in the browser toolbar does look up genes for you; that lookup is not available to the agent.)
Known app defects
The defects below were recorded in earlier testing. Unless a note says otherwise, they have not been re-tested against the current code, so some may already be fixed.
- Peak BED tracks do not render. Tracks written by
genome_call_peaksshow the message "Error detecting chromosome naming" instead of intervals. The BED file content is fine: Genie writes a plain, unindexed BED file and registers it as abedtrack with a URL, and the browser engine reads such URLs with an indexed (tabix) reader. The BED files in the project's results folder can be opened elsewhere. (The current code still writes peak results as plain BED tracks of typebed; rendering was not re-tested for this page.) - First-load drawing defect. On first load the browser panel draws only about 72% of the labelled region, cutting off the right end. The first zoom, or a window resize, clears it.
- Share dialog fails in local mode. View-only share links exist only in cloud mode; in the local backend the share dialog shows "Sharing settings could not be loaded." (The current local backend still has no sharing route.)
- Raw tokens in saved Reports. A Studio Report saved in Analysis mode can contain a raw, unresolved citation token in its HTML.
- Viewport statistics cited to region-only records. As described above, the citation after viewport statistics opens only the locus. (This is still the case in the current code.)
- Failure results still get a citation token. When a computation tool returns a failure (for example, a scope over 5 Mb), the failure is still cited with a token whose excerpt is the failure, and the call is reported to the runtime as successful, so its row shows as done.
- Other display issues. Typing a locus into the toolbar Gene symbol box does not navigate, and tables in the chat can break words mid-word in a narrow chat panel.
For how sharing works in cloud mode, see Export and sharing.