Skip to main content

Public data search

public_data_search is the Genome-mode tool that finds public data for you. It searches the ENCODE portal only. The agent turns your request into structured fields, the backend queries the portal, and the tool returns a ranked list of released files with their accessions, a download URL and a citation token for each. Loading a file is a separate step: the agent calls genome_add_track with the chosen candidate's URL and type.

The Genome-mode instructions tell the agent to fill the fields from your words itself, and not to claim that a track represents a tissue, assay, target or accession without copying that candidate's citation token after the claim.

Fields​

FieldDefaultWhat it does
assayENCODE assay title, for example "Histone ChIP-seq", "ATAC-seq", "DNase-seq"
targetTarget label, for example "H3K27ac" or "CTCF"
biosampleBiosample term, for example "K562"; matched exactly first, by organ or cell slim only in a relaxation
organismHomo sapiensScientific name of the donor organism
assemblythe project genomehg38 is sent as GRCh38 and hg19 as hg19, whatever the field says; for other genomes the field, or else the genome id, is sent
fileTypebigWigFile format (bigWig, bigBed or another ENCODE format)
freeTextSent as the portal's free-text search term
limit10Number of candidates returned, clamped to 1–25

Every portal request asks for released records only and times out after 12 seconds.

Illustratively, a request such as "Find public H3K27ac ChIP-seq fold-change signal for K562 on hg38" gives the agent what it needs for assay, target and biosample; the assembly comes from the project and the file type defaults to bigWig.

Two stages​

  1. Experiments. If you gave an assay, target, biosample or free text, the tool first searches ENCODE experiments with those filters (assay title, target label, biosample term, organism, assembly, status released). It asks for up to 3 × limit experiments, and at least 20.
  2. Files. It then searches files of the requested format and assembly that belong to the experiments it found, asking for up to 4 × limit files. If the experiment stage found no experiments, the file search applies the assay, target, biosample, organism and free-text filters directly.

The first query and the relaxation ladder​

The first file query also asks for output type "fold change over control". If a step returns no candidates, the next step runs, up to three relaxations:

StepRuns whenChangeRecorded as
0alwaysOutput type "fold change over control", biosample by exact term(no relaxation)
1step 0 returned nothingDrop the output-type filter"dropped output_type=fold change over control"
2step 1 returned nothing and a biosample was givenRedo the experiment search, matching the biosample by organ slim"used biosample_ontology.organ_slims"
3step 2 returned nothing and a biosample was givenRedo the experiment search, matching the biosample by cell slim"used biosample_ontology.cell_slims"

The relaxations taken are returned in relaxations, and every portal query, with its URL and counts, in attempts. Each candidate also carries a matchMode (term, organ_slim or cell_slim).

Check for relaxations

A relaxed search can return a different kind of file from the one you asked for: a different output type after step 1, or a related biosample after steps 2 and 3. The agent is not forced to tell you. Expand the tool card and look at relaxations and at each candidate's outputType and biosample.

Example of a relaxation (GATA1 test session). ENCODE has no DNase-seq files of output type "fold change over control", so the search for a K562 DNase-seq signal track returned nothing at step 0 and found a file only after the first relaxation. The loaded file, ENCFF008EKC, is read-depth normalized signal. The agent named the DNase track's output type but did not report the relaxation, or flag comparability when it later reported a correlation with the H3K27ac fold-change track.

Ranking​

Each candidate gets a fixed additive score:

ConditionPoints
Status released+20
Two or more biological replicates+10
Output type contains "fold change", "signal", "read depth" or "normalized"+8
Biosample matches your term exactly+8
Biosample given and matched only through an organ or cell slim+3
Target matches exactly+6
Download URL uses https+2

Duplicates of the same file accession are removed, keeping the higher score. Candidates are sorted by score and then by accession, alphabetically, and cut to limit. No date is used: the search does not prefer newer files. Because all queries ask for released records, every candidate gets the +20.

In the MYC and ALB demonstrations, the K562 H3K27ac search ranked ENCFF381NDD first through the accession tie-break among candidates with equal scores.

What a candidate contains​

FieldMeaning
rank, titlePosition and a label such as "H3K27ac · K562 · GRCh38 · ENCFF381NDD"
accession, experimentENCODE file (ENCFF...) and experiment (ENCSR...) accessions
assay, target, biosample, assemblyFrom the file record, or from its experiment
fileType, outputType, biologicalReplicatesAs recorded by ENCODE
sourceUrlDownload URL (the file's cloud URL, or the portal download link); passed to genome_add_track
typeTrack type for genome_add_track (bigwig, bigbed, bedgraph)
portalUrlhttps://www.encodeproject.org/files/<accession>/
citationTokenToken for claims about this file

The search also returns a top-level citation token for claims about the result set as a whole.

Each search stores one public-data-search record (the normalized query, the number of candidates returned and the relaxations) and one public-data-file record per candidate. A file record holds the accession, experiment, assay, target, biosample, assembly and output type; its Open source link opens the ENCODE file page. In the MYC demonstration, chip [2] opened the record for ENCFF381NDD. See Citations and evidence.

Loading a result​

The search does not add tracks. The agent loads a candidate with genome_add_track, passing sourceUrl, type and a name. The added track keeps only its id, type, name and URL; the assay and output type travel with it only if the agent puts them in the name.

Live portal​

The tool queries the live ENCODE portal. Its results, and their order, can change as ENCODE adds, revises or archives records, so the same request can return different candidates on another day. Searches are not cached. If you need to repeat an analysis, record the accession and load it directly.

How this differs from Public Data Hubs​

The browser's Public Data Hubs sheet (File → Open Public Data Hubs…) is a separate, manual route. It lists curated hub collections for the project's genome, such as ENCODE and Roadmap hubs, from hub lists hosted at vizhub.wustl.edu, and you choose and add tracks yourself.

public_data_search (agent)Public Data Hubs (UI)
Who drives itThe agent, from your requestYou
SourceLive ENCODE portal searchCurated hub lists
RankingFixed score, then accessionNone; you browse and filter
Citation recordsYes, per search and per fileNo
Adds tracksNo; a separate genome_add_track callYes, when you click +

The agent cannot load a hub: genome_add_track rejects hubUrl as not supported yet. See Adding tracks.