Public data search
public_data_search is the Genome-mode tool that finds public data for you. It searches the ENCODE portal only. The agent turns your request into structured fields, the backend queries the portal, and the tool returns a ranked list of released files with their accessions, a download URL and a citation token for each. Loading a file is a separate step: the agent calls genome_add_track with the chosen candidate's URL and type.
The Genome-mode instructions tell the agent to fill the fields from your words itself, and not to claim that a track represents a tissue, assay, target or accession without copying that candidate's citation token after the claim.
Fields
| Field | Default | What it does |
|---|---|---|
assay | ENCODE assay title, for example "Histone ChIP-seq", "ATAC-seq", "DNase-seq" | |
target | Target label, for example "H3K27ac" or "CTCF" | |
biosample | Biosample term, for example "K562"; matched exactly first, by organ or cell slim only in a relaxation | |
organism | Homo sapiens | Scientific name of the donor organism |
assembly | the project genome | hg38 is sent as GRCh38 and hg19 as hg19, whatever the field says; for other genomes the field, or else the genome id, is sent |
fileType | bigWig | File format (bigWig, bigBed or another ENCODE format) |
freeText | Sent as the portal's free-text search term | |
limit | 10 | Number of candidates returned, clamped to 1–25 |
Every portal request asks for released records only and times out after 12 seconds.
Illustratively, a request such as "Find public H3K27ac ChIP-seq fold-change signal for K562 on hg38" gives the agent what it needs for assay, target and biosample; the assembly comes from the project and the file type defaults to bigWig.
Two stages
- Experiments. If you gave an assay, target, biosample or free text, the tool first searches ENCODE experiments with those filters (assay title, target label, biosample term, organism, assembly, status released). It asks for up to 3 ×
limitexperiments, and at least 20. - Files. It then searches files of the requested format and assembly that belong to the experiments it found, asking for up to 4 ×
limitfiles. If the experiment stage found no experiments, the file search applies the assay, target, biosample, organism and free-text filters directly.
The first query and the relaxation ladder
The first file query also asks for output type "fold change over control". If a step returns no candidates, the next step runs, up to three relaxations:
| Step | Runs when | Change | Recorded as |
|---|---|---|---|
| 0 | always | Output type "fold change over control", biosample by exact term | (no relaxation) |
| 1 | step 0 returned nothing | Drop the output-type filter | "dropped output_type=fold change over control" |
| 2 | step 1 returned nothing and a biosample was given | Redo the experiment search, matching the biosample by organ slim | "used biosample_ontology.organ_slims" |
| 3 | step 2 returned nothing and a biosample was given | Redo the experiment search, matching the biosample by cell slim | "used biosample_ontology.cell_slims" |
The relaxations taken are returned in relaxations, and every portal query, with its URL and counts, in attempts. Each candidate also carries a matchMode (term, organ_slim or cell_slim).
A relaxed search can return a different kind of file from the one you asked for: a different output type after step 1, or a related biosample after steps 2 and 3. The agent is not forced to tell you. Expand the tool card and look at relaxations and at each candidate's outputType and biosample.
Example of a relaxation (GATA1 test session). ENCODE has no DNase-seq files of output type "fold change over control", so the search for a K562 DNase-seq signal track returned nothing at step 0 and found a file only after the first relaxation. The loaded file, ENCFF008EKC, is read-depth normalized signal. The agent named the DNase track's output type but did not report the relaxation, or flag comparability when it later reported a correlation with the H3K27ac fold-change track.
Ranking
Each candidate gets a fixed additive score:
| Condition | Points |
|---|---|
| Status released | +20 |
| Two or more biological replicates | +10 |
| Output type contains "fold change", "signal", "read depth" or "normalized" | +8 |
| Biosample matches your term exactly | +8 |
| Biosample given and matched only through an organ or cell slim | +3 |
| Target matches exactly | +6 |
| Download URL uses https | +2 |
Duplicates of the same file accession are removed, keeping the higher score. Candidates are sorted by score and then by accession, alphabetically, and cut to limit. No date is used: the search does not prefer newer files. Because all queries ask for released records, every candidate gets the +20.
In the MYC and ALB demonstrations, the K562 H3K27ac search ranked ENCFF381NDD first through the accession tie-break among candidates with equal scores.
What a candidate contains
| Field | Meaning |
|---|---|
rank, title | Position and a label such as "H3K27ac · K562 · GRCh38 · ENCFF381NDD" |
accession, experiment | ENCODE file (ENCFF...) and experiment (ENCSR...) accessions |
assay, target, biosample, assembly | From the file record, or from its experiment |
fileType, outputType, biologicalReplicates | As recorded by ENCODE |
sourceUrl | Download URL (the file's cloud URL, or the portal download link); passed to genome_add_track |
type | Track type for genome_add_track (bigwig, bigbed, bedgraph) |
portalUrl | https://www.encodeproject.org/files/<accession>/ |
citationToken | Token for claims about this file |
The search also returns a top-level citation token for claims about the result set as a whole.
Citations from a search
Each search stores one public-data-search record (the normalized query, the number of candidates returned and the relaxations) and one public-data-file record per candidate. A file record holds the accession, experiment, assay, target, biosample, assembly and output type; its Open source link opens the ENCODE file page. In the MYC demonstration, chip [2] opened the record for ENCFF381NDD. See Citations and evidence.
Loading a result
The search does not add tracks. The agent loads a candidate with genome_add_track, passing sourceUrl, type and a name. The added track keeps only its id, type, name and URL; the assay and output type travel with it only if the agent puts them in the name.
Live portal
The tool queries the live ENCODE portal. Its results, and their order, can change as ENCODE adds, revises or archives records, so the same request can return different candidates on another day. Searches are not cached. If you need to repeat an analysis, record the accession and load it directly.
How this differs from Public Data Hubs
The browser's Public Data Hubs sheet (File → Open Public Data Hubs…) is a separate, manual route. It lists curated hub collections for the project's genome, such as ENCODE and Roadmap hubs, from hub lists hosted at vizhub.wustl.edu, and you choose and add tracks yourself.
public_data_search (agent) | Public Data Hubs (UI) | |
|---|---|---|
| Who drives it | The agent, from your request | You |
| Source | Live ENCODE portal search | Curated hub lists |
| Ranking | Fixed score, then accession | None; you browse and filter |
| Citation records | Yes, per search and per file | No |
| Adds tracks | No; a separate genome_add_track call | Yes, when you click + |
The agent cannot load a hub: genome_add_track rejects hubUrl as not supported yet. See Adding tracks.