For any research lab with published or open data

Validated computational models built from your lab's own data.

OMEGA is a computational discovery system. It reads a lab's published or openly licensed data through a connector, builds models validated on held-out data, and sends candidates back for the lab to measure. What comes back from the lab — confirmed, refuted, or inconclusive — is recorded against the candidate that produced it.

  • 1  No fabricated data: every number comes from a real record or a stated computation.
  • 2  Held-out validation before a model is trusted; never a random split reported as a result.
  • 3  Every claim carries what would refute it, and a refutation is recorded as readily as a confirmation.

What is connected today

Every figure below is rendered by the server from the live system when this page is served. None of them is typed into the page.

system digest written 2026-09-13T16:15+00:00

18
labs connected through a real connector
64
datasets those connectors list
33,026
real normalized records read from those labs
0
schema violations in the last read (shown even when zero)
78
candidates prepared for delivery back to labs
12
packages drafted and waiting for a person to send them
0
packages a person has sent
0
lab results received so far. Zero means no lab has answered yet, and it is shown as zero.
63
attempts the system recorded as failed, each with the gate that refused it and the reason
39,862
candidates the system has adopted from its own reviewed queue
32
research domains those adoptions span

A candidate prepared or a package drafted is an output, not an outcome. The only outcome this system counts is a result a lab measured and sent back; that count is the one above labelled results received.

How it works

Six steps. Two of them are a person's, and the system does not take those two by itself.

1

Register

Your lab's name, a contact email, and your institution. No account, no password.

2

Receive an API key

Shown once, at registration. Only a hash of it is stored, so it cannot be shown again.

3

Submit a dataset in the record contract

Each record is validated against the same contract every hand-built connector meets. Violations come back per record, by index.

4

A person reviews it

A valid submission is kept on file and waits for a human decision. Nothing is wired in automatically, however clean it validates.

5

A connector is built

Written by hand against your real data, and read on the same schedule and under the same checks as every other connected lab.

6

Candidates come back with a reply token

A person emails you a package. Each carries a token and states what result would refute it. You answer at /reply, and your answer is recorded against that package.

WHAT WE ASK FOR

Data you already have, with its provenance

  • Published or openly licensed measurements, one record per measured thing, in the record contract shown on the connect page.
  • Where each record came from: a non-empty source, a verification status, and for records taken from an open repository the license and a record id or DOI.
  • Your lab's stated research goal in a sentence, so candidates are ranked against what you are trying to do rather than against what happens to be easy to compute.
  • Optionally, unpublished measurements under whatever agreement your group needs. That is the one thing no literature search can reach, and it is never required.
  • A reply when a package arrives. A refutation is as useful as a confirmation, and both are recorded.
WHAT YOU GET

Ranked candidates, and the record of what failed

  • A connector that reads your data under the same contract, checks, and schedule as every other connected lab, with its record count shown on this page and its schema violations counted in the total above.
  • Candidates ranked by a model validated on held-out data, each stating the measurement that would refute it.
  • The failures too: what the system tried against your data and could not do, with the gate that refused it and why.
  • Nothing published or acted on without your review. The system computes; it does not synthesize, measure, or send.
  • What it does not give you: a wet-lab result. Every candidate is a computation until your lab measures it, and this site counts it that way.

Connected labs

Each row is a connector as the system last read it. Institution and department are what the connector states; the goal is the lab's own stated research aim.

Connector Institution Department Stated goal Datasets Records Last read
nci60_nci_cellminercdb National Cancer Institute Predict which drugs a given cancer cell line will respond to via a structurally independent panel (NCI's own 60-line collection, distinct compound library and cell lines from GDSC/CTRP) -- the same applied purpose, a third real cross-check. 2 25,929 2026-09-09
doench_broad Broad Institute Predict CRISPR-Cas9 sgRNA on-target editing efficiency reliably enough to design a working guide for any genomic target without empirical trial-and-error -- the practical purpose Rule Set 2 exists to serve, not merely 'fit this one dataset'. 4 3,913 2026-09-09
mit_miller_lab Massachusetts Institute of Technology Brain and Cognitive Sciences / Picower Institute for Learning and Memory Our research examines the neural basis of cognition and consciousness, including executive functions like working memory, attention, and decision-making. We study how the brain learns abstract rules, categories, and concepts. We study how neural dynamics, like oscillatory brain rhythms, contribute to cognition and consciousness. Our aim is to understand cognition and consciousness to develop treatments for autism, schizophrenia, and attention deficit disorder, and to improve anesthesia safety (ekmillerlab.mit.edu/research, read directly 2026-09-07). 9 1,046 2026-09-09
gdsc_sanger_mgh Wellcome Sanger Institute Predict which drugs a given cancer cell line/tumour genotype will actually respond to, and find the genomic biomarkers that explain why -- prioritizing real compounds for further testing and matching patients to therapies, the applied purpose behind the dose-response measurements, not the IC50/AUC numbers as an end in themselves. 2 400 2026-09-09
gtex_v10_v11 GTEx Consortium Understand tissue-specific gene expression/regulation well enough to identify disease-relevant targets and assess on-target/off-tumour safety risk BEFORE a therapeutic candidate reaches the clinic -- the applied purpose GTEx feeds elsewhere in this codebase (omega_cure/antigen_specificity.py's CAR-T safety gate), not description for its own sake. 2 400 2026-09-09
mit_jazayeri_lab Massachusetts Institute of Technology Brain and Cognitive Sciences / McGovern Institute for Brain Research Understand how dynamic patterns of neural activity in the brain (as well as in artificial neural systems) give rise to algorithms that support mental computations -- reverse-engineer the brain's algorithms by seeing the brain as a system of recurrent neural networks generating complex neural dynamics, and develop a mathematical framework connecting these levels of description, working with humans, animal models and recurrent neural network models (jazlab.org, read directly 2026-09-09). 9 231 2026-09-09
ctrp_broad_ctd2 Broad Institute CTD2 Center Predict which drugs a given cancer cell line will respond to and find the real gene targets behind that response -- the same applied purpose as GDSC's own end goal, from a structurally independent screen (different compound library, different institution). 1 200 2026-09-09
lincs_l1000_broad Broad Institute Infer a candidate drug's real mechanism of action and find repurposing candidates via connectivity mapping -- does a compound's own measured transcriptomic signature reverse a disease's real expression signature -- the applied purpose behind the perturbation signatures, distinct from viability screening (GDSC/CTRP/NCI-60/PRISM already cover that axis; this is a genuinely different, complementary readout of the same real compounds). 1 200 2026-09-09
matbench_materials_project Materials Project Discover and rank NEW, not-yet-synthesized stable crystal structures via ML-predicted formation energy -- the benchmark exists to validate models for real materials discovery, not to re-fit the 132,752 already-known Materials Project entries it ships with. 1 200 2026-09-09
aprahamian_dartmouth Dartmouth College Chemistry Design and validate new hydrazone-based molecular switches/motors that convert controlled molecular motion into mechanical work (drug delivery, energy storage); new hydrazone-based fluorophores for bio-imaging, sensing, and OLED applications; and new hydrazone-based sensing systems for chemical/biological detection -- the group's own 3 stated research pillars (aprahamiangroup.com/research.html, verified live 2026-08-11), not just the switching-kinetics subset this connector's existing data happens to cover. 4 152 2026-09-09
t_agency_self_study OMEGA (this project) T-AGENCY Understand how human/LLM collaborative scientific creativity, intuition, and innovation actually manifest as a process -- structurally, with real checkable evidence, not narrative -- well enough to deliberately improve the loop itself, not just document what happened after the fact. Unlike this connector's 5 wet-lab/public-dataset siblings, the relevant 'other labs' here are cognitive-science/HCI/computational-creativity research groups studying human-AI collaboration in general, not a domain-adjacent chemistry/biology dataset. 20 149 2026-09-09
prism_repurposing_broad Broad Institute Predict which (often already-approved) drugs a given cancer cell line will respond to, the repurposing angle specifically -- surfacing existing drugs for a new cancer indication rather than discovering a wholly new compound. 1 112 2026-09-09
turk_dartmouth_nccc Dartmouth College Norris Cotton Cancer Center Generate durable memory T-cell responses to cancer -- since tumors are altered self-tissue rather than foreign pathogens, lasting immune memory against them does not form the way it does against pathogens. The lab's own stated findings: autoimmune disease is a critical determinant for maintaining T-cell memory to tumor antigens; the most functional tumor-specific memory T cells reside in peripheral tissues, not blood; tissue-resident memory (TRM) T cells are a critical component of long-lived anti-tumor immunity (geiselmed.dartmouth.edu faculty page, read directly 2026-08-22). 1 32 2026-09-09
rosenberg_nci_surgery_branch National Cancer Institute Surgery Branch Develop adoptive cell transfer immunotherapy -- isolating and expanding a patient's own tumor-reactive T cells (TIL) ex vivo, then reinfusing them with IL-2 support -- to durably regress metastatic cancer; pioneered at the NCI Surgery Branch (Steven A. Rosenberg) and extended by real methodological advances (e.g. NeoExpand neoantigen-specific stimulation) that isolate more, and functionally better (stem-like memory), tumor-reactive TCRs than conventional screening. 1 25 2026-09-09
disc_varshney_gaurav_oklahoma_medical_researc OKLAHOMA MEDICAL RESEARCH FOUNDATION Development of Precision Genome Editing Tools for the Functional Characterization of Genetic Variants 1 24 2026-09-11
disc_engreitz_jesse_stanford_university STANFORD UNIVERSITY An integrative genomics toolkit for rewriting regulatory DNA to reprogram gene expression 3 6 2026-09-11
disc_bauer_daniel_boston_children_s_hospit BOSTON CHILDREN'S HOSPITAL Computational tools for precision genome editing 1 5 2026-09-11
disc_musunuru_kiran_children_s_hosp_of_phila CHILDREN'S HOSP OF PHILADELPHIA Postnatal and Prenatal Therapeutic Base Editing for Metabolic Diseases 1 2 2026-09-11

Contact

Questions about the contract, a data-use agreement, attribution, or a lost API key: write to us. A person reads it.

[email protected]