M2iOR data collection is comprised in 2 different tables:
- Experiment Table: Gathers all the OR-molecules experiments, with detailed description of both molecule and receptor as well as the response and the assay realized.
- Blast Table: Aggregated view on the OR sequences using BLAST algorithm.
The Experiment Table can be divided into five parts: Receptor and Co-receptor Information, Molecule Description, Experimental Response, Bioassay Description, and Resources.
Receptor and Co-Receptor:
- Species: Taxon name of the species from which the OR originates.
- Receptor/Co-Receptor Name: Gene Name of the given OR/Orco.
- Accession: Identifier in the UniprotKB or Genbank database.
- Mutation: Mutation realized by the authors on a given protein sequence.1
- Sequence: Protein sequence of the given OR.2
1 For mutated receptors/co-receptors, the wild-type identifier and sequence is retrieved, and the mutation is separately indicated using XpositionY format where amino acid X from the wild-type sequence is mutated to amino acid Y at the given position. Additions are indicated by setting X to "ins" and deletion with "del" as Y.
2 Protein sequence serves as a standardized OR identifier in M2iOR. If it is not available in the original publication, it is retrieved from Uniprot or Genbank using the name or other identifier provided by the authors.
Molecule:
- Molecule Name: Name of the molecule.1
- CID: Compound Identifier from PubChem.2
- CAS: CAS Registry Number.1
- InChIKey: International Chemistry Identifier Key from PubChem.2,3
- SMILES: Simplified Molecular-Input Line-Entry System, which is a text representation of the structure of the molecule.3,4
- Mixture: Distinguish between blend of multiple molecules (“mixture”), combination of isomers (“sum of isomers”), or a mono molecular compound (“mono molecular”).5
1 Since CAS is a proprietary system and chemical nomenclature is not fully standardized, the compound name or CAS number stored in M2iOR may differ slightly from the terminology used in the original publication. In such cases, the recorded values correspond to PubChem's canonical entry for the resolved compound (compound name and CAS synonym) rather than to the wording used by the authors.
2 As part of the standardized curation pipeline of M2iOR, the PubChem Compound ID (CID) serves as the primary molecular identifier of stimuli. A molecular entry is included only if the curated information obtained from the source publication reports either a CID or a CAS registry number. When a CID is provided, it is used to query PubChem, which returns the compound name, CAS number, SMILES, and International Chemical Identifier Key (InChIKey). These annotations replace the values reported in the original publication whenever discrepancies are found. If the publication reports only a CAS number, the CAS identifier is first mapped to a CID using a PubChem name-based search. The corresponding compound name, CAS number, SMILES, and InChIKey are then retrieved from the resolved PubChem entry.
3 Mixtures are identified by a space-separated list for both InChIKey and SMILES, for each of its components.
4 If only the molecular structure is available, a SMILES representation was inferred and further used to search for the corresponding InChIKey.
5 The diversity of the composition of the tested compounds is defined by three classes: ”mixture”, ”sum of isomers” and ”mono molecular”. The term ”mixture” indicates that the experiments were carried out using a blend of multiple molecules (e.g. essential oils or artificial composition) and mixtures are identified by a space-separated list of InChIKeys of each compound. For all the molecules, information on isomerism is carefully researched. Number of chiral centers and geometric isomerism are automatically determined from SMILES using the RDKit package. If a molecule has at least one chiral center and no specific information is provided about the enantiomery or the diasteroisomery, the molecule is a combination of isomers, and thus labelled ”sum of isomers”. In cases where the authors explicitly mentioned isomerism of a tested molecule, it is identified as ”mono molecular”. All achiral molecules are also labeled as ”mono molecular”.
Response:
- Responsive: Binary-encoded response of the pair, with FALSE representing a non-agonist and TRUE an agonist1.
- Parameter: The type of response obtained in the experiment and stored in the database: ”primary” for a single dose, ”secondary” for 2–3 concentrations, and ”dose-response” for a full dose-response curve.
- Value: The value reported by the source (if applicable) or EC50 value for dose response measurements2.
- Unit: Unit of the measurement reported in Value column. (firing rate, potential, no unit in case of normalized response, …).
- Value Nature: Type of value reported in Value column, can be ”EC50” for dose-response measurement, ”Raw” if the raw response of the receptor is available or “Norm_rec”, “Norm_pair”, “Norm_other” for responses normalized by either the receptor’s baseline, response of a given pair or an unknown normalization denominator 2.
- Concentration: The concentration used for a given screening experiment (“Raw“ data) or the concentration range (if available) used to determined “EC50“ in a dose-response experiment.
- Concentration Unit: Unit of the value reported in Concentration column.
- Nbr. Measurement: Number of repetitions for a given experiment.
1 Decision about the response in the Responsive column, are solely made by the respective authors. If authors did not provide conclusion on the responsiveness a specific workflow is used to determine the response.
2 Value can be “n.d.” for non-responsive pairs determined in dose-response measurements. When a mixture is tested, when available, the concentration of each of its components is described in a space-separated list.
Bioassay:
- Experimental technique: The type experimental technique used to make measurement (e.g., “Two-electrode voltage clamp“, “calcium imaging“, “single sensillum recording“, ...)
- Recording system: The recording system used to make the expertiments (e.g., “robot model for oocytes“, “manual two-electrode voltage clamp“, “tungstene electrode SSR“, “glass electrode SSR“, “calcium dyes for calcium imaging“, …)
- Type: The measured quantity which could be a concentration of “Ca2+”, “spiking frequency“, “current amplitude“, ”fluorescence”…
- Expression System: The type of cell or organism used to make the experiments (e.g., ”HEK”, ”Xenopus oocytes”, ”Drosophila”, …)
- Expression Cell Type: The specific cell type used for heterologous expression within the expression system, when applicable (e.g., ”HEK293”, ”ab3a”, ”at1/ab3”, …).
- Expression Driver: The genetic driver or promoter used to control receptor expression in the model organism, if relevant (e.g., ”Orco-GAL4”, ”OR67d-GAL4”, ”pCS2+”, ”pT7Ts”, …)
- Co-Transfection: Any protein co-transfected with the olfactory receptors, as mentioned in the source.
- Odor delivery system: The method or device used to deliver odorants to the preparation (e.g., ”liquid”, ”airborne paper”, ”airborne headspace”, ”airborne GC”, …).
- Stimulation Flux: The flow rate of the air/liquid stream carrying the odorant during stimulation.
- Stimulation Unit: The unit used for Stimulation Flux (e.g., ”mL/s”, ”L/min”, …).
- Main Flux: The background or carrier flow rate of the main air/liquid stream into which the odorant is injected. Relevant for systems with constant background flow.
- Main Flux Unit: The unit used for Main Flux.
- Stimulation Duration: The duration of odorant presentation for a given stimulus in the experiment.
- Stimulation Duration Unit: The unit used for Stimulation Duration (e.g., ”s”, ”ms”).
Resources:
- Reference: Bibliographic reference of the source.
- DOI: Digital Object Identifier (DOI) of the source.
- Reference Position: Specific location of the information about the pair in the source. (e.g., Table1, Fig S1...)
The insect genome includes a large family of odorant receptor genes, known as IRs, ORs, and GRs depending on the lineage. In Drosophila melanogaster, for example, there are approximately 60 canonical odorant receptor (OR) genes, with many species showing lineage-specific expansions or losses. Unlike vertebrate ORs, insect ORs are ligand-gated ion channels formed by a heteromeric complex of a conserved co-receptor (Orco) and a variable odorant-binding OR subunit. Each OR subunit has a distinct recognition spectrum and alterations in one or more amino acids can significantly change its response.
Multiple variants and mutants of the same gene have been tested in the literature and they are gathered in the M2iOR database. Some of these sequences show different responses compared to the reference. To facilitate the comparison of such cases, similar sequences are grouped under the same reference sequence.
BLAST algorithm is used to compare each sequence in M2iOR against the Uniprot (2026 release 02) or Genbank (release 271, April 15 2026) database. The Uniprot/Genbank ID of the best match in terms of identity are then associated with these sequences. They are subsequently grouped by their best match’s name, resulting in the BLAST table.
- Receptor Name: Receptor Name of the given OR.
- Accession: Identifier in the Uniprot/Genbank database.
- % Seq Identity: Percentage of identity between the sequence in Experiment Table and the sequence in Uniprot/Genbank.
- Species: Taxon name of the species from which the OR originates.
- Mutation1: Differences in amino acids between the sequence in Uniprot and the Sequence in M2iOR.
- Sequence: Protein sequence of the given OR in Experiment Table.
1 The mutation is indicated using XpositionY format where amino acid X from the Uniprot/Genbank sequence is mutated to amino acid Y at the given position. Additions are indicated by setting X to "ins" and deletion with "del" as Y.