Computation profile for Mechanistic Coherence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Null-Variant Quality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of null-variant quality.
bemo
BEMO:2000226
BEMO:2000226
233
de2f4749e64a7014070a28c9c8587efaf4dbb9b3583c2ab6e78cd6826c00de34
8
Assesses null-variant quality using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in null-variant quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating null-variant quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Systems-Level Emergence Support
The magnitude and credibility of independent evidence supporting systems-level emergence.
bemo
BEMO:2000024
BEMO:2000024
31
07df8f6070c405006e36806e8ccb515126dc2c7269ea9cf3934662d28e9264c6
8
Assesses systems-level emergence support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in systems-level emergence support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating systems-level emergence support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Transcript Quantification Reliability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of transcript quantification reliability.
bemo
BEMO:2000265
BEMO:2000265
272
45c762be688f8b3f57f3b563b58fc816d5802e731de9aee9acb732eea429ea1a
8
Assesses transcript quantification reliability using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in transcript quantification reliability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating transcript quantification reliability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Prespecified Analysis Adherence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prespecified analysis adherence.
bemo
BEMO:2000423
BEMO:2000423
430
fe370510831c4feb18d187a9f8348ee0b93ff3174da9ddceea906467255bbd60
8
Assesses prespecified analysis adherence using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in prespecified analysis adherence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating prespecified analysis adherence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Mechanistic Coverage
The proportion and representativeness of the relevant mechanistic captured by the evidence or measurement process.
bemo
BEMO:2000012
BEMO:2000012
19
8582aac321a237f1dd1db68f0893021d3bc987b71b373d1879cc2dfb098adbdb
9
Assesses mechanistic coverage using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in mechanistic coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating mechanistic coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Evidence Sufficiency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Derivatization Efficiency
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of derivatization efficiency.
bemo
BEMO:2000360
BEMO:2000360
367
d83b78fc8beda508a5c18da8dbb9e284c948a9741d4fb8d070f9f30bc9267827
8
Assesses derivatization efficiency using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in derivatization efficiency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating derivatization efficiency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 106
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Dechallenge–Rechallenge Support; Effect-Modification Credibility; Collider Bias Risk", "Common Misinterpretations": "Treating mediation evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Mediation Evidence Strength", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The magnitude and credibility of independent evidence supporting mediation evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses mediation evidence strength using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in mediation evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
106
b7abd5f34cd2addb34f26d923a9c80b402d8914f57510a7f5fe8c4ed41659c58
Computation profile for Incorporation Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Call-Rate Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Information Size Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Bioavailability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of bioavailability.
bemo
BEMO:2000331
BEMO:2000331
338
1dbec5a00f53d17cad5a3a88dd4dfde9ceab9bbbe33a898c76c7086698d1a0a2
8
Assesses bioavailability using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in bioavailability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating bioavailability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Knowledge-Graph Evidence Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Developing; BEMO computation profile requires independent validation.
Computation profile for Selectivity Profile
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Adverse Outcome Pathway Support
The magnitude and credibility of independent evidence supporting adverse outcome pathway.
bemo
BEMO:2000329
BEMO:2000329
336
8935703e7b30761e4213b86ddafc1aecae0696f4523a201860d9e64493990a74
8
Assesses adverse outcome pathway support using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in adverse outcome pathway support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating adverse outcome pathway support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Cell-Type Annotation Confidence
The justified degree of certainty assigned to cell-type annotation given the quantity, quality, consistency, and limitations of supporting evidence.
bemo
BEMO:2000241
BEMO:2000241
248
f7aebc612affb172059739df32608fb7c4d69384f1ff8c565432595eea6b9937
10
Assesses cell-type annotation confidence using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in cell-type annotation confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cell-type annotation confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 377
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Metabolite Identification Confidence; Spectral Library Match Quality; Internal Standard Performance", "Common Misinterpretations": "Treating metabolite annotation level as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Metabolite Annotation Level", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of metabolite annotation level.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses metabolite annotation level using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in metabolite annotation level can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
377
883e5fdd2ff9e198b3ddcbafc3e6478eed144bbcc7f46c67f19ec93850e8ec4c
Minimal Clinically Important Difference Validity
The degree to which minimal clinically important difference supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000047
BEMO:2000047
54
604e83a95f413c6c8ddc8abe7ad03f24c441ede8e65d506baef0410866c23fa7
10
Assesses minimal clinically important difference validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in minimal clinically important difference validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating minimal clinically important difference validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Control Group Appropriateness
The extent to which control group is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000473
BEMO:2000473
480
47dcb7d2162b3c8496d326c0cd9b9dfbef0e0b4032c2c7d93f31f67139c5990f
8
Assesses control group appropriateness using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in control group appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating control group appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Missing-Data Reporting Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Safety Biomarker Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Test-Timing Appropriateness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Biological Plausibility
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biological plausibility.
bemo
BEMO:2000002
BEMO:2000002
9
6cbc5f525d668cebb0928b8428848a625103a7b6492d0a2a618adb116f200bd3
9
Assesses biological plausibility using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in biological plausibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biological plausibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Selective Outcome Reporting Risk
The probability or degree that selective outcome reporting introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000486
BEMO:2000486
493
bcdaa5f06df0a5ab7941c490a740787093ffc20f433c48125d9c2e174850f7fd
8
Assesses selective outcome reporting risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in selective outcome reporting risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating selective outcome reporting risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Module Stability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Evidence Stability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 285
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Reference Interval Validity; Measurement Uncertainty; Method Comparison Agreement", "Common Misinterpretations": "Treating cutoff validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Cutoff Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The degree to which cutoff supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses cutoff validity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cutoff validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
285
f74e484a95fd67b9bd73e0f0c112bc85ddaf01e7476f82b9c24c9382fe2bfe48
Computation profile for Post-Translational Modification Localization Confidence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Noninferiority Margin Validity
The degree to which noninferiority margin supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000452
BEMO:2000452
459
5c494db875d0f3ad6860a3bba83e6ac01d9f571b40aa89617a37a42c1c489cab
10
Assesses noninferiority margin validity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in noninferiority margin validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating noninferiority margin validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Genotype Quality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of genotype quality.
bemo
BEMO:2000247
BEMO:2000247
254
6b3ff41c46f655393bf02a8730b778b41b8fad3c6f62302d7543aef197484ea3
8
Assesses genotype quality using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in genotype quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.
Treating genotype quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Temporal Generalizability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Knowledge-Graph Provenance Quality
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Developing; BEMO computation profile requires independent validation.
Perturbation Prediction Accuracy
The closeness of perturbation prediction to the accepted reference or true value.
bemo
BEMO:2000324
BEMO:2000324
331
5cc466b3de703eb697570667c57a9230690c56d69b5cf6ec528e42740ec621e9
8
Assesses perturbation prediction accuracy using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in perturbation prediction accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating perturbation prediction accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 315
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Knowledge-Graph Provenance Quality; Relation Evidence Strength; Ontology Annotation Completeness", "Common Misinterpretations": "Treating entity resolution accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Entity Resolution Accuracy", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The closeness of entity resolution to the accepted reference or true value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses entity resolution accuracy using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in entity resolution accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
315
81ae3d3ac205c8d1258be22f4b5818a5c36df354b19605060eb458cfdfef682e
Evidence Stability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence stability.
bemo
BEMO:2000153
BEMO:2000153
160
1f9ca5bc5be56489e2caf9d2f3eaaf37b12ae4271716853fb373d2db6aa8d657
9
Assesses evidence stability using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Between-Study Variance
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of between-study variance.
bemo
BEMO:2000139
BEMO:2000139
146
167d33dabeef5837833710d7e97d42ff17e46ede74b98e0e88db46e921ff811e
8
Assesses between-study variance using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in between-study variance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Forest plots; heterogeneity statistics; tau-squared; prediction intervals; funnel plots; regression or selection models; sensitivity analyses.
Treating between-study variance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Time–Concentration Profile Adequacy
0.1.0
Apply a validated analyte- and method-specific measurement procedure with calibration and quality control.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","unitRef":"REQUIRED"}
specimen_or_material; measurement_procedure; calibration_reference; quality_control_results; unit
replicate_measurements; environmental_conditions; instrument_version
xsd:decimal
Method- and analyte-specific physical units
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Derivatization Efficiency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Instrument Drift
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Technical Replicate Adequacy
The extent to which technical replicate is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000189
BEMO:2000189
196
8f5bd5eed372632eafdff28f1c8035067d4634c207012fb089ff55a67d705176
8
Assesses technical replicate adequacy using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in technical replicate adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating technical replicate adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Diagnostic Odds Ratio
0.1.0
value = numerator / denominator; the null value is typically 1 where scientifically applicable
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator != 0"],"null_value":1}
numerator; denominator; operational_definition; assessment_context
confidence_level; stratum
xsd:decimal
Ratio scale; null typically 1
0.0
1
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Not generally required unless converted to a probability or score.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Variant Pathogenicity Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Reproducibility of Measurement
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Cross-Omics Integration Coherence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cross-omics integration coherence.
bemo
BEMO:2000304
BEMO:2000304
311
c10b406fc4ce6db5663280fc4b358cc062cc17f0ce7a593af68a762307815cdc
8
Assesses cross-omics integration coherence using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in cross-omics integration coherence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cross-omics integration coherence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Developing
Computation profile for Peptide-Spectrum Match Quality
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 138
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Reclassification Improvement; Prognostic Calibration; Calibration Slope", "Common Misinterpretations": "Treating prognostic discrimination as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Prognostic Discrimination", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prognostic discrimination.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses prognostic discrimination using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in prognostic discrimination can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
138
b58deaf6b24f17f45025c2e10bd742dcb00a937cc0410366fbe092bfbfd13882
Processed-Data Availability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of processed-data availability.
bemo
BEMO:2000424
BEMO:2000424
431
3d4a65e6d3c66bcaca4ff40b6df33feaedf2f77c58697c32f80c96a00d211e95
8
Assesses processed-data availability using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in processed-data availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating processed-data availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Duplicate Read Burden
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of duplicate read burden.
bemo
BEMO:2000245
BEMO:2000245
252
c1d6b63aea87c65750f96b5e3a771c6866843a2e5ba7e8081dd07af03bc0e9c5
8
Assesses duplicate read burden using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in duplicate read burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.
Treating duplicate read burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Biospecimen Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Instrument Validity
The degree to which instrument supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000098
BEMO:2000098
105
e7033f91a59f115f9fba748b04cab66664ac151f4c8d653ea61aa823a8c6cce4
10
Assesses instrument validity using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in instrument validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating instrument validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Decision-Curve Net Benefit
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of decision-curve net benefit.
bemo
BEMO:2000436
BEMO:2000436
443
7e96f3c494e389efab04f0ca564863ebd840cadeaf424446d3f94ddb483ed76c
8
Assesses decision-curve net benefit using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in decision-curve net benefit can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating decision-curve net benefit as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Time-to-Fixation Adequacy
The extent to which time-to-fixation is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000080
BEMO:2000080
87
49559a7ae628351af963a27a55893cf501098be34876ce35bdc2d4969fd4f94b
8
Assesses time-to-fixation adequacy using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in time-to-fixation adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating time-to-fixation adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Experimental Unit Validity
The degree to which experimental unit supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000176
BEMO:2000176
183
3d08eaaab51eb4d1095b9fb5e884110ba36de86e61615acc1661c581f985a6b1
10
Assesses experimental unit validity using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in experimental unit validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating experimental unit validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Isotope Pattern Fidelity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of isotope pattern fidelity.
bemo
BEMO:2000367
BEMO:2000367
374
beec5d62727ea3baba10451d532801205b190af60b03027284c15989d51ba1e4
8
Assesses isotope pattern fidelity using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in isotope pattern fidelity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating isotope pattern fidelity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 299
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Preanalytical Robustness", "Common Misinterpretations": "Treating postanalytical integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Postanalytical Integrity", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of postanalytical integrity.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses postanalytical integrity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in postanalytical integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
299
b1cf2cbf0a36602510f89d4d570874b9d2148d8bd132ef3b47e7be72655f9ff3
Computation profile for Penetrance Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Rule Based Rubric Computation
Controlled BEMO ComputationMode value: RuleBasedRubricComputation.
BEMO:4000011
RuleBasedRubricComputation
Recruitment Reporting Completeness
The extent to which all scientifically necessary components of recruitment reporting are present, documented, and evaluable.
bemo
BEMO:2000429
BEMO:2000429
436
d9e587b4cd42324cc1668fb835e848ce02fea72bac0af2f37bcc6e0a40cbc327
9
Assesses recruitment reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in recruitment reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating recruitment reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Biomarker Clinical Relevance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Loss-of-Function Mechanism Validity
The degree to which loss-of-function mechanism supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000225
BEMO:2000225
232
a389d5411c414282451fe5495bf2c1c444f06ae12cd17457edf93df2953dc32e
10
Assesses loss-of-function mechanism validity using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in loss-of-function mechanism validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating loss-of-function mechanism validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Posterior Probability Strength
The magnitude and credibility of independent evidence supporting posterior probability.
bemo
BEMO:2000456
BEMO:2000456
463
8a96194c8b8e65d04df266b282704e480656791b819df12c8c4924c984b02272
8
Assesses posterior probability strength using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in posterior probability strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating posterior probability strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Response Biomarker Validity
The degree to which response biomarker supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000053
BEMO:2000053
60
a8e8ef21a701fbb547c028557e4678bc79ebeda610eb6926bf0061b6e9a016c4
10
Assesses response biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in response biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating response biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Receptor Occupancy Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pathology Confirmation
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of pathology confirmation.
bemo
BEMO:2000074
BEMO:2000074
81
e470a75855d7bb24a7f3957922ccdf78961ef4c923b6f754c0006ebb46c0367e
8
Assesses pathology confirmation using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in pathology confirmation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating pathology confirmation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Trial-Level Surrogacy
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of trial-level surrogacy.
bemo
BEMO:2000057
BEMO:2000057
64
977d5f2b7e02b5e90f442ea4e8788e2fe539316cddda3ae61da29b0fa71c5325
8
Assesses trial-level surrogacy using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in trial-level surrogacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating trial-level surrogacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 477
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Contamination Risk; Carryover Effect Risk; Period Effect Risk", "Common Misinterpretations": "Treating co-intervention bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Co-intervention Bias Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that co-intervention bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses co-intervention bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in co-intervention bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
477
de36d6b209ddc70bf44fd968006b0ba61482bcc4bbe42db3c2a605301631de5e
Causal Identifiability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of causal identifiability.
bemo
BEMO:2000086
BEMO:2000086
93
a6559bfe5ff78dd7455ed538b35cec6508c7c5507441bc614a71e2140188258e
10
Assesses causal identifiability using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in causal identifiability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.
Treating causal identifiability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 424
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Materials-and-Reagents Reporting Completeness; Raw-Data Availability; Processed-Data Availability", "Common Misinterpretations": "Treating metadata completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Metadata Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of metadata are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses metadata completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in metadata completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
424
cfce2b3ca58f108e1dc5461ad1ef2f5a84b55171c75536d306bf4e1e2b3bc58d
Computation profile for Independent Replication Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Case-Control Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Adherence Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Assumption Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of assumption sensitivity.
bemo
BEMO:2000084
BEMO:2000084
91
b75d62759fbd8f126f27997543f0f46dde1cb8c2484d156a08c9925c9476aa63
9
Assesses assumption sensitivity using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in assumption sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating assumption sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Analytical Validity
The degree to which analytical supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000272
BEMO:2000272
279
e4b1deae92fd2c549054f6f19f234dbf0b73ab8ff30ff5741db1c88b5618f6c3
10
Assesses analytical validity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in analytical validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating analytical validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Parameter Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of parameter sensitivity.
bemo
BEMO:2000321
BEMO:2000321
328
9735c06335e80bafcbeea74ccb7b405e18e8684330b82448424f6e83ef3974ba
9
Assesses parameter sensitivity using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in parameter sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating parameter sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Housing and Husbandry Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Protein Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 471
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Randomization Integrity; Baseline Comparability; Blinding Integrity", "Common Misinterpretations": "Treating allocation concealment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Allocation Concealment", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of allocation concealment.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses allocation concealment using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in allocation concealment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
471
3367e7e58461dcb460260b263a504d906f8c26af8608d0cf2767790d3843a2e7
Computation profile for Relatedness Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Cell-Type Specificity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Batch-Effect Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of batch-effect sensitivity.
bemo
BEMO:2000273
BEMO:2000273
280
7222c30713cc035f86bae406549aaf0d4751a919d50ae68b917efd9629244d28
9
Assesses batch-effect sensitivity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in batch-effect sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating batch-effect sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Computational Variant Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 495
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Control Group Appropriateness; Study Design Appropriateness; Protocol Fidelity", "Common Misinterpretations": "Treating temporal precedence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Temporal Precedence", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of temporal precedence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses temporal precedence using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in temporal precedence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
495
ec85a4abfc9e480fea03dfff5177f17a293899d488f06a679715de258c8591b9
uncertainty required
Whether uncertainty reporting is mandatory.
BEMO:3100016
Comparator Applicability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of comparator applicability.
bemo
BEMO:2000192
BEMO:2000192
199
607cceee390b7cb8ce577becf5db87c7c9a4643da6c8fbe5bd44948e1b9e3853
8
Assesses comparator applicability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in comparator applicability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating comparator applicability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Spectral Library Match Quality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of spectral library match quality.
bemo
BEMO:2000385
BEMO:2000385
392
5f1fa5732761f5ddf9eae216997fca90d2ec95264368ca476160849aac179e9a
8
Assesses spectral library match quality using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in spectral library match quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating spectral library match quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Reclassification Improvement
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 191
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Sample Size Justification; Blinding in Experimental Assessment; Animal Model Face Validity", "Common Misinterpretations": "Treating randomization in experimental allocation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Randomization in Experimental Allocation", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of randomization in experimental allocation.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses randomization in experimental allocation using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in randomization in experimental allocation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
191
5cc94b6085b9c596d4f74c18712e9f8177db4335ab703b2218b9c3a63209c113
Animal Model Construct Validity
The degree to which animal model construct supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000167
BEMO:2000167
174
f8c928bcc279a4856d6377104c402fb4a00d102300c3c3fa14181f174b53f974
10
Assesses animal model construct validity using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in animal model construct validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating animal model construct validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Phenocopy Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Hook Effect Risk
The probability or degree that hook effect introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000281
BEMO:2000281
288
3d6a1547d736b7a958bf1c94a2df8cb5599400f10305a91c657ae1dce184650b
8
Assesses hook effect risk using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in hook effect risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating hook effect risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Quality-Control Reporting Completeness
The extent to which all scientifically necessary components of quality-control reporting are present, documented, and evaluable.
bemo
BEMO:2000427
BEMO:2000427
434
b53b586f994fd48c27b9acc1c6fff59b966085610e593ec05c03735d90e329b2
9
Assesses quality-control reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in quality-control reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating quality-control reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 359
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Therapeutic Window Evidence; No-Observed-Adverse-Effect Level Robustness; Lowest-Observed-Adverse-Effect Level Robustness", "Common Misinterpretations": "Treating safety margin evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Safety Margin Evidence", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of safety margin evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses safety margin evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in safety margin evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
359
51ce0324d6952becbbadfd71ba2913a1f56afbcc33aa0148093d10994d18697d
Computation profile for Clinical Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Generic Template Defined
Controlled BEMO FormulaStatus value: GenericTemplateDefined.
BEMO:4000019
GenericTemplateDefined
Anatomical Site Fidelity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of anatomical site fidelity.
bemo
BEMO:2000058
BEMO:2000058
65
8fa3b411fb892fff8cc31c078e25770eb74980a525cbc8e76d63bcb0f15bdaf8
8
Assesses anatomical site fidelity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in anatomical site fidelity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating anatomical site fidelity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 169
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Missing Evidence Risk; Selective Nonreporting Risk; Small-Study Effects", "Common Misinterpretations": "Treating publication bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Publication Bias Risk", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The probability or degree that publication bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Forest plots; heterogeneity statistics; tau-squared; prediction intervals; funnel plots; regression or selection models; sensitivity analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses publication bias risk using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in publication bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
169
ae14a6a3126f2ecd31144006a017263d5e10df0fcde0e6bd8cc0872e93853a7a
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 267
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Alternative Splicing Validation; Single-Cell Ambient RNA Burden; Single-Cell Viability", "Common Misinterpretations": "Treating single-cell doublet burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Single-Cell Doublet Burden", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of single-cell doublet burden.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses single-cell doublet burden using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in single-cell doublet burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
267
47e6d4bd3f64b35d8abb9925b05e112fbb1356d383dc641a47f78f8e9bbeb676
Storage Temperature Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of storage temperature control.
bemo
BEMO:2000079
BEMO:2000079
86
b39a92acfa55bd840331823c88476d9d65ca0179f68f0f37ed8d7677a63d2193
8
Assesses storage temperature control using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in storage temperature control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating storage temperature control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Genome Coverage Uniformity
The proportion and representativeness of the relevant genome uniformity captured by the evidence or measurement process.
bemo
BEMO:2000246
BEMO:2000246
253
c86e8ba8cbfb79ea4dbc1a9a4cd5da76837e35c7f45f5966b51b9734760bdf1f
9
Assesses genome coverage uniformity using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in genome coverage uniformity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating genome coverage uniformity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 132
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Area Under the Receiver Operating Characteristic Curve; Threshold Validity; Reference Standard Validity", "Common Misinterpretations": "Treating partial area under the receiver operating characteristic curve as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Partial Area Under the Receiver Operating Characteristic Curve", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of partial area under the receiver operating characteristic curve.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses partial area under the receiver operating characteristic curve using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in partial area under the receiver operating characteristic curve can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
132
79f6b607e0d7504f46e1d3842fa8966ae929db7058236a6b4ed553277ad444b5
Multiplicity-Adjusted Credibility
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of multiplicity-adjusted credibility.
bemo
BEMO:2000159
BEMO:2000159
166
40900cc3965adf1a4a61ca1880158e0ed27ede5d1a6163f63eb4f6d38b99a3b8
8
Assesses multiplicity-adjusted credibility using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in multiplicity-adjusted credibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating multiplicity-adjusted credibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Adherence Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of adherence integrity.
bemo
BEMO:2000463
BEMO:2000463
470
9a20a1494ccbc46d7bc0acd4d260241ebc29edd15b77d8ccc525c9f0dba8e8de
8
Assesses adherence integrity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in adherence integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating adherence integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Outcome Definition Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Lipemia Burden
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of lipemia burden.
bemo
BEMO:2000070
BEMO:2000070
77
572ca1cf50802a5447be4c1286cdad2930a7ec08e97e7b9af663a3cb66e1d788
8
Assesses lipemia burden using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in lipemia burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating lipemia burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Susceptibility/Risk Biomarker Validity
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Cross-Species Biological Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Missing-Data Mechanism Plausibility
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of missing-data mechanism plausibility.
bemo
BEMO:2000447
BEMO:2000447
454
399e0c1d410c7ca1c5cc767b5afb9291915d1ee3db3b4d7314b2092c6d4765e1
9
Assesses missing-data mechanism plausibility using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in missing-data mechanism plausibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating missing-data mechanism plausibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Warm Ischemia Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of warm ischemia control.
bemo
BEMO:2000083
BEMO:2000083
90
293c5ebda3d3695852f06644549a5b920266b83358bf49b383b5d90078fcdde6
8
Assesses warm ischemia control using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in warm ischemia control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating warm ischemia control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Statistical Methods Reporting Completeness
The extent to which all scientifically necessary components of statistical methods reporting are present, documented, and evaluable.
bemo
BEMO:2000431
BEMO:2000431
438
d8af14618f3e80ecff3b617a35c54caa9b3c73184f234fcea88a9ee83c7090ff
9
Assesses statistical methods reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in statistical methods reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating statistical methods reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Reference Interval Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 88
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Freeze–Thaw Burden; Processing Delay Control; Anatomical Site Fidelity", "Common Misinterpretations": "Treating transport condition integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Transport Condition Integrity", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of transport condition integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses transport condition integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in transport condition integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
88
81f1c97831f3ffbcc8b4e89c16604f987d3fa2baac750496d710981c68a60598
Safety Margin Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of safety margin evidence.
bemo
BEMO:2000352
BEMO:2000352
359
51ce0324d6952becbbadfd71ba2913a1f56afbcc33aa0148093d10994d18697d
8
Assesses safety margin evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in safety margin evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating safety margin evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Environmental Standardization
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of environmental standardization.
bemo
BEMO:2000173
BEMO:2000173
180
432e96e9d5c02fcda1a353b0ddb2e18d1019ff594f33c248eedfbe337bf31fae
8
Assesses environmental standardization using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in environmental standardization can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating environmental standardization as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Mass Accuracy
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Entity Resolution Accuracy
The closeness of entity resolution to the accepted reference or true value.
bemo
BEMO:2000308
BEMO:2000308
315
81ae3d3ac205c8d1258be22f4b5818a5c36df354b19605060eb458cfdfef682e
8
Assesses entity resolution accuracy using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in entity resolution accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating entity resolution accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Diagnostic Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of diagnostic sensitivity.
bemo
BEMO:2000116
BEMO:2000116
123
4ea75714eea731a1e4644a00db0476f8628bf028934f92bd9e4029ee31ddf173
9
Assesses diagnostic sensitivity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in diagnostic sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.
Treating diagnostic sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Pharmacological Target Validity
The degree to which pharmacological target supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000348
BEMO:2000348
355
2cff4752043ea892788f37639d922b1a249d83c13dfedf581ad1a4ad12ab95b8
10
Assesses pharmacological target validity using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in pharmacological target validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating pharmacological target validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 228
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Variant Pathogenicity Evidence Strength; Population Frequency Compatibility", "Common Misinterpretations": "Treating gene–disease validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Gene–Disease Validity", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The degree to which gene–disease supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses gene–disease validity using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in gene–disease validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
228
4501a0d637c0d3d7b3852d3be8f61115b2b1ac442582dd9a3ec88bd21dad9f7c
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 408
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Data Provenance Completeness; Material Availability; Code Availability", "Common Misinterpretations": "Treating protocol reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Protocol Reproducibility", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The degree to which protocol reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses protocol reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in protocol reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
408
caa28491bc96ceb6cd2e1c6685b9eb89f21391a6d4502473143986af13cc088b
Functional Variant Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of functional variant evidence.
bemo
BEMO:2000220
BEMO:2000220
227
7a71bec12dfb580d6d32593ad450868d51a755f2a8b35e7eee350ad98d24c73e
8
Assesses functional variant evidence using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in functional variant evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating functional variant evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Dilution Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of dilution integrity.
bemo
BEMO:2000279
BEMO:2000279
286
a31ef0b724223ebc3d1ad0dbe731b0b7823f6b347840ea43769f01a33866757b
8
Assesses dilution integrity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in dilution integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating dilution integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Sex as a Biological Variable Adequacy
The extent to which sex as a biological variable is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000186
BEMO:2000186
193
67d1bc84f2c2d0324ee734ed3d2df1c72cd4820b21290b5a8b136d75863a4866
8
Assesses sex as a biological variable adequacy using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in sex as a biological variable adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating sex as a biological variable adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Founder-Effect Assessment
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Evidence Quality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence quality.
bemo
BEMO:2000151
BEMO:2000151
158
f4e08b1965ce7d1832ff4c5191f1e7298781d9bec3cb9a96c09198d6faaa1339
8
Assesses evidence quality using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Organ-Specific Toxicity Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 369
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Internal Standard Performance; Derivatization Efficiency; Metabolic Feature Reproducibility", "Common Misinterpretations": "Treating extraction recovery as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Extraction Recovery", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of extraction recovery.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses extraction recovery using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in extraction recovery can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
369
5024c9034e231e8bc296cf1d9092cc063fe94890f4856d8bf70ac077514b55f9
Computation profile for Carryover Effect Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Dynamic Range
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of dynamic range.
bemo
BEMO:2000280
BEMO:2000280
287
41309e5f36c397d08d76621c273137352b099c93484f2a2eca025bcf85ccffdc
8
Assesses dynamic range using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in dynamic range can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating dynamic range as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Under Review
Controlled BEMO ApprovalStatus value: UnderReview.
BEMO:4000028
UnderReview
Missing-Data Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of missing-data sensitivity.
bemo
BEMO:2000448
BEMO:2000448
455
6708217d617153345c000ff0b34cc96bd34facf9514e5d1a84b3e56623ffb545
9
Assesses missing-data sensitivity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in missing-data sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating missing-data sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Reproducibility of Measurement
The degree to which reproducibility of measurement yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000299
BEMO:2000299
306
83e10528c5f3255388bb807d6bfc99c4e2c93d4f418421fcd184c242521982d5
10
Assesses reproducibility of measurement using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in reproducibility of measurement can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reproducibility of measurement as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 304
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Analytical Precision; Intermediate Precision; Reproducibility of Measurement", "Common Misinterpretations": "Treating repeatability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Repeatability", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The degree to which repeatability yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses repeatability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in repeatability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
304
289ee6101d32dc64bb4644380815480f0d1de4b3d0512c355bf6d1bb8e30c7f2
Pathway Enrichment Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of pathway enrichment robustness.
bemo
BEMO:2000373
BEMO:2000373
380
1a891a794290e0f5eee9ca88b64bc6b46d5840f63c8709e0e546399158fbb914
8
Assesses pathway enrichment robustness using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in pathway enrichment robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating pathway enrichment robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Comparator Applicability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 486
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Exposure Classification Validity; Comparator Validity; Control Group Appropriateness", "Common Misinterpretations": "Treating intervention classification validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Intervention Classification Validity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The degree to which intervention classification supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses intervention classification validity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in intervention classification validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
486
283a3544a9dd5cb1a16190c929c82c6db2608db0f4c24860780498cb89bc0ac5
Reference Interval Validity
The degree to which reference interval supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000296
BEMO:2000296
303
3706e41c2a6c519ccdfb2643facf47ec1d2a45cc4b5b902e54040a5a53624b7e
10
Assesses reference interval validity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in reference interval validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reference interval validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Measurement Error Correction
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of measurement error correction.
bemo
BEMO:2000446
BEMO:2000446
453
47dddb5f2dde48130ebfa9464b01b7ed1efb5da34ac2d484caa4259055162e55
8
Assesses measurement error correction using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in measurement error correction can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating measurement error correction as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Materials-and-Reagents Reporting Completeness
The extent to which all scientifically necessary components of materials-and-reagents reporting are present, documented, and evaluable.
bemo
BEMO:2000416
BEMO:2000416
423
fc72795f141f7cb81bc8cc86190a0c7cf2a5f2f6065d3b22f9bc88dd2e13080e
9
Assesses materials-and-reagents reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in materials-and-reagents reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating materials-and-reagents reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Evidence Consensus Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for External Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Outlier Influence Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of outlier influence robustness.
bemo
BEMO:2000454
BEMO:2000454
461
3a52b33f3faa79b95100ff084b9c7ca54158523c656fcaf5c298f073960428c9
8
Assesses outlier influence robustness using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in outlier influence robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating outlier influence robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Experimental Batch Randomization
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of experimental batch randomization.
bemo
BEMO:2000175
BEMO:2000175
182
1291027c6f53e5902a86a2cf22b3cca02564dbc940590bb8b8bcfd5004b998b7
10
Assesses experimental batch randomization using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in experimental batch randomization can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating experimental batch randomization as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Pathology Confirmation
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Attrition Accounting in Animal Studies
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of attrition accounting in animal studies.
bemo
BEMO:2000170
BEMO:2000170
177
bfb6735bf1e0494f94802ec087edc274ccdaea50558f3ea043f6f21ae9210505
8
Assesses attrition accounting in animal studies using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in attrition accounting in animal studies can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating attrition accounting in animal studies as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Threshold Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Decision-Curve Net Benefit
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Care-Pathway Independence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of care-pathway independence.
bemo
BEMO:2000191
BEMO:2000191
198
e39c9004563f895f591d754f0c2c01d5cfe45a0e26d44584ecb91c07d72c494a
8
Assesses care-pathway independence using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in care-pathway independence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating care-pathway independence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Prospective Registration
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Mechanistic Coherence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of mechanistic coherence.
bemo
BEMO:2000010
BEMO:2000010
17
dc1b3b375698d0aec42693d5a0b2f87405e7e7bf90759105797ad714ab4e9ab2
8
Assesses mechanistic coherence using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in mechanistic coherence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating mechanistic coherence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Criterion Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Discriminant Validity
The degree to which discriminant supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000041
BEMO:2000041
48
c0e22c7d45b8b17a4c9f1de5c7c7bb4e085d98eb667da8d43e452ea31955900b
10
Assesses discriminant validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in discriminant validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating discriminant validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 453
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Overadjustment Bias Risk; Calibration of Statistical Predictions; Discrimination of Statistical Predictions", "Common Misinterpretations": "Treating measurement error correction as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Measurement Error Correction", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of measurement error correction.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses measurement error correction using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in measurement error correction can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
453
47dddb5f2dde48130ebfa9464b01b7ed1efb5da34ac2d484caa4259055162e55
Computation profile for Data-Sharing Transparency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Evidence Consistency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Quantitative Bias Analysis Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of quantitative bias analysis robustness.
bemo
BEMO:2000103
BEMO:2000103
110
7ee1bd197cfd8211ed7ac785b0e49ec8e73324814e415213789e479477f73c74
10
Assesses quantitative bias analysis robustness using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in quantitative bias analysis robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating quantitative bias analysis robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Fixation Adequacy
The extent to which fixation is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000066
BEMO:2000066
73
c08e61688a36fec680e9073fc8888dbfd98852919e3f7ccf7c3180ed49d1e921
8
Assesses fixation adequacy using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in fixation adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating fixation adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Molecular-Phenotypic Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Contamination Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Allelic Evidence Strength
The magnitude and credibility of independent evidence supporting allelic evidence.
bemo
BEMO:2000211
BEMO:2000211
218
2dd4a685750034c6175101f0128d889784cf621fed0cf93a74098e6b4879332b
8
Assesses allelic evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in allelic evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating allelic evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Method Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Relatedness Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of relatedness control.
bemo
BEMO:2000255
BEMO:2000255
262
c66c714b8c2b1fb9da08eabc61ab1b346b5733232f7901baf90ed1206b3491ba
8
Assesses relatedness control using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in relatedness control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating relatedness control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Single-Cell Feature Detection Rate
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of single-cell feature detection rate.
bemo
BEMO:2000261
BEMO:2000261
268
1ce87bfa0e14b3c49cdf26e92f3b39ba4ce32c6192d394a99b4f3b80c80031b2
8
Assesses single-cell feature detection rate using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in single-cell feature detection rate can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating single-cell feature detection rate as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Target Trial Emulation Fidelity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Dynamic Range Coverage
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Latent-Factor Stability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of latent-factor stability.
bemo
BEMO:2000312
BEMO:2000312
319
be1e726b2fd3735d5492ec8f375d473b36cbff860abc6d682d2065974c115c50
9
Assesses latent-factor stability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in latent-factor stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating latent-factor stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Statistical Methods Reporting Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Outcome Ascertainment Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Exclusion-Criteria Prespecification
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Exposure Classification Validity
The degree to which exposure classification supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000477
BEMO:2000477
484
6d08552ad6c8b3c6d5a3c8ed35293fcfc624984197116fb158f6c3a04698789f
10
Assesses exposure classification validity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in exposure classification validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating exposure classification validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Prognostic Biomarker Validity
The degree to which prognostic biomarker supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000052
BEMO:2000052
59
2dd804ff14e01e089d83a5d2d6ef790d24cf90ca54d3426049116e4e7aeb095f
10
Assesses prognostic biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in prognostic biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating prognostic biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Experimental Batch Randomization
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Differential Follow-up Risk
The probability or degree that differential follow-up introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000475
BEMO:2000475
482
ccfe7e445aeb97f5f6fbabfec00706b7cd315705f68e05e908e47860b40d8ad4
8
Assesses differential follow-up risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in differential follow-up risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating differential follow-up risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 127
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Comparative Test Accuracy; Reclassification Improvement; Prognostic Discrimination", "Common Misinterpretations": "Treating incremental diagnostic value as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Incremental Diagnostic Value", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of incremental diagnostic value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses incremental diagnostic value using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in incremental diagnostic value can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
127
8c0d8dc6983b3100a525f45a2c3c327af7d21cf2d155dc05a19b99e0d7f43578
Hemolysis Burden
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of hemolysis burden.
bemo
BEMO:2000068
BEMO:2000068
75
1b0946c6b109834d9e6fa6dc5ae7dde1ce95bc65f999a58a3c5a88bc66c6a47c
8
Assesses hemolysis burden using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in hemolysis burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating hemolysis burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Confidence Interval Compatibility
The justified degree of certainty assigned to confidence interval compatibility given the quantity, quality, consistency, and limitations of supporting evidence.
bemo
BEMO:2000435
BEMO:2000435
442
4b44e65362db0ef8f1ef0ae6eb87b631fd22642d5790f792ab6c3d6a9a693508
10
Assesses confidence interval compatibility using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in confidence interval compatibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating confidence interval compatibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Immunotoxicity Evidence Strength
The magnitude and credibility of independent evidence supporting immunotoxicity evidence.
bemo
BEMO:2000340
BEMO:2000340
347
577a4370c7a1dc270882fc655889fdc7a0b8ee20d5e6f81f1f733233db939fbd
8
Assesses immunotoxicity evidence strength using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in immunotoxicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating immunotoxicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Segregation Evidence Strength
The magnitude and credibility of independent evidence supporting segregation evidence.
bemo
BEMO:2000232
BEMO:2000232
239
4f9bdcfda63260ced3f40796d43228eea03ead68f517bc87939698d4f5d4faf2
8
Assesses segregation evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in segregation evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating segregation evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Preservation Adequacy
The extent to which preservation is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000075
BEMO:2000075
82
de3923232eba7ea0cab4f8eddc6a1539eac9f06416d7c31e6d2ce8fbf8492ba5
8
Assesses preservation adequacy using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in preservation adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating preservation adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Method Reproducibility
The degree to which method reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000399
BEMO:2000399
406
b36c4ba0356ae6813b5bb255ebc6934635b0cbf83c44326cfea7019592803c5e
10
Assesses method reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in method reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating method reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Model–Experiment Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Fragility Index
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Protocol Reproducibility
The degree to which protocol reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000401
BEMO:2000401
408
caa28491bc96ceb6cd2e1c6685b9eb89f21391a6d4502473143986af13cc088b
10
Assesses protocol reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in protocol reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating protocol reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Effect Magnitude
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 332
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Structural Identifiability; Dynamical Stability; Steady-State Validity", "Common Misinterpretations": "Treating practical identifiability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Practical Identifiability", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of practical identifiability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses practical identifiability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in practical identifiability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
332
cb110edeb57b36960477192ea5d1484823905d046fc0f22c33bc8eb340c3c4c3
Computation profile for Competing-Risk Model Validity
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Evidence Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence robustness.
bemo
BEMO:2000152
BEMO:2000152
159
59315ce160586081fc3fd71f96646716829d26fde9e0a21aa0afec923b5083d6
8
Assesses evidence robustness using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Biospecimen Provenance Completeness
The extent to which all scientifically necessary components of biospecimen provenance are present, documented, and evaluable.
bemo
BEMO:2000060
BEMO:2000060
67
c859051c1fe5f3cb60cda0bb4dcf804459be135ee18bb1bb44d89345dea50253
9
Assesses biospecimen provenance completeness using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in biospecimen provenance completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biospecimen provenance completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Drug–Drug Interaction Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Spectrum Representativeness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of spectrum representativeness.
bemo
BEMO:2000207
BEMO:2000207
214
771c727a328004aa438f83d20410cd3600ce67ac5affa57fc6e9aaee55cb7ed9
8
Assesses spectrum representativeness using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in spectrum representativeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating spectrum representativeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Network Node Confidence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Postanalytical Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of postanalytical integrity.
bemo
BEMO:2000292
BEMO:2000292
299
b1cf2cbf0a36602510f89d4d570874b9d2148d8bd132ef3b47e7be72655f9ff3
10
Assesses postanalytical integrity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in postanalytical integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating postanalytical integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 240
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Null-Variant Quality; RNA Evidence Strength; Co-segregation Likelihood", "Common Misinterpretations": "Treating splicing evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Splicing Evidence Strength", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The magnitude and credibility of independent evidence supporting splicing evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses splicing evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in splicing evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
240
50c6384810a1443b5a98f42466fb8ab2489b57dad29ed2dab35ef7acec1960d2
Computation profile for Construct Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Expressivity Consistency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Randomization Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Biomarker Responsiveness
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Flux-Balance Consistency
The degree of agreement in flux-balance across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000309
BEMO:2000309
316
f8d3b19efd68e0150e437c71bac90ab2bb80f33877b4a0e74bc52542a0fd3c26
8
Assesses flux-balance consistency using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in flux-balance consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating flux-balance consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Calibration Traceability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of calibration traceability.
bemo
BEMO:2000274
BEMO:2000274
281
7d6efe540a48fdb6c1d86b381d867ad016fd03d3ba3039e34ba0dd8d59bcbf8d
8
Assesses calibration traceability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in calibration traceability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating calibration traceability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Metabolite Coverage
The proportion and representativeness of the relevant metabolite captured by the evidence or measurement process.
bemo
BEMO:2000342
BEMO:2000342
349
4bb05358971c228c81cd2238f505b14da043cabfc3f68e3d2dd07884c7f72bc5
9
Assesses metabolite coverage using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in metabolite coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating metabolite coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Cross-Layer Directional Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
scientific claim
An information content entity asserting a biomedical proposition that may be evaluated using evidence metrics.
bemo
BEMO:0000202
Candidate
Human-Relevance of Toxicological Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of human-relevance of toxicological evidence.
bemo
BEMO:2000339
BEMO:2000339
346
32b4aadfeeab004a7a018b1fe09f04f0b2283eb92e7749123f8ac828db8cfce6
8
Assesses human-relevance of toxicological evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in human-relevance of toxicological evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating human-relevance of toxicological evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Known-Groups Validity
The degree to which known-groups supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000046
BEMO:2000046
53
8e195ef346deba393b7a3adaf4c77bfbd949dbf0f057cd9af26312568cb85421
10
Assesses known-groups validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in known-groups validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating known-groups validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Small-Study Effects
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of small-study effects.
bemo
BEMO:2000164
BEMO:2000164
171
ec5ee159f748d5035209ae456b71092c6724bc73af9b720bf70afe3ffc41f242
8
Assesses small-study effects using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in small-study effects can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Forest plots; heterogeneity statistics; tau-squared; prediction intervals; funnel plots; regression or selection models; sensitivity analyses.
Treating small-study effects as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Cell-Type Annotation Confidence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Carcinogenicity Evidence Strength
The magnitude and credibility of independent evidence supporting carcinogenicity evidence.
bemo
BEMO:2000332
BEMO:2000332
339
a171459b41fd9ad0436588a80669a48227118c77c3d87194654b56ab6dfa6ad7
8
Assesses carcinogenicity evidence strength using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in carcinogenicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating carcinogenicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Random-Seed Stability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of random-seed stability.
bemo
BEMO:2000402
BEMO:2000402
409
c1a5a2e956afa3ff1606fb1893282ee7b8d5d3733fbd04798630fe0746f86074
9
Assesses random-seed stability using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in random-seed stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating random-seed stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Positivity Adequacy
The extent to which positivity is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000102
BEMO:2000102
109
be2ac5c0345d954f0ccc1ce41f05b792425f3fb9c158083a6db1928102b69608
8
Assesses positivity adequacy using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in positivity adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.
Treating positivity adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Baseline Comparability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Convergent Validity
The degree to which convergent supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000038
BEMO:2000038
45
4ad4866845885f677cad74042f65b0303f2d26597ea694b76dfed4bc71fedd68
10
Assesses convergent validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in convergent validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating convergent validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Comparator Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Early Stopping Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Selective Nonreporting Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Prognostic Added Value
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prognostic added value.
bemo
BEMO:2000129
BEMO:2000129
136
6a8e54c8d71bbaa2bc8361d6fc19eb1f50b8f4acd2b6d09e2b648f2337fa9298
8
Assesses prognostic added value using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in prognostic added value can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating prognostic added value as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Cold Ischemia Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cold ischemia control.
bemo
BEMO:2000063
BEMO:2000063
70
450f9e94b009036abe7a13fe3055116e7c4959857a7291ed7e337a746d7bf722
8
Assesses cold ischemia control using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in cold ischemia control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cold ischemia control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Technical Artifact Exclusion
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Raw-Data Availability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of raw-data availability.
bemo
BEMO:2000428
BEMO:2000428
435
2fcf4f96e7ca202f61d3de161ead38ea506d30ac8f05b66ac8a84b27f3da645e
8
Assesses raw-data availability using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in raw-data availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating raw-data availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Founder-Effect Assessment
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of founder-effect assessment.
bemo
BEMO:2000219
BEMO:2000219
226
6530e53b02059e4bfd6465ae64ac385198a560d8a80df4ba701894a8c5cb0975
8
Assesses founder-effect assessment using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in founder-effect assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating founder-effect assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Rescue Experiment Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Attrition Accounting in Animal Studies
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Strand Bias
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of strand bias.
bemo
BEMO:2000264
BEMO:2000264
271
6ac08afc8cf99d9cd5ee94bc247ff32edf260fc542fd5e27409d9c3b702afd66
10
Assesses strand bias using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in strand bias can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.
Treating strand bias as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Contamination Burden
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Analytical Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Protocol Availability
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 392
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Metabolite Annotation Level; Internal Standard Performance; Extraction Recovery", "Common Misinterpretations": "Treating spectral library match quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Spectral Library Match Quality", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of spectral library match quality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses spectral library match quality using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in spectral library match quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
392
5f1fa5732761f5ddf9eae216997fca90d2ec95264368ca476160849aac179e9a
Computation profile for Positive-Control Performance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
formula status
A controlled concept describing whether a formula is defined, generic, method-specific, or context-specific.
bemo
BEMO:0000403
Candidate
Computation profile for Minimal Clinically Important Difference Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Humane Endpoint Appropriateness
The extent to which humane endpoint is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000179
BEMO:2000179
186
efafe26c06dac15c0d485bfb264933c40fe9cf7ca97c5195ea92c9cfeb3f7a98
8
Assesses humane endpoint appropriateness using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in humane endpoint appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating humane endpoint appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Overall Evidence Certainty
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of overall evidence certainty.
bemo
BEMO:2000160
BEMO:2000160
167
31b48110148a1ebdd1379016bfff93898087ec61573fac3b850ccada3e1f7ae5
10
Assesses overall evidence certainty using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in overall evidence certainty can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating overall evidence certainty as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
No-Observed-Adverse-Effect Level Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of no-observed-adverse-effect level robustness.
bemo
BEMO:2000344
BEMO:2000344
351
f5f71661574e809d9eb1b7625ffde603665b4c0d25ddbb968f6b8d1a88c38a42
8
Assesses no-observed-adverse-effect level robustness using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in no-observed-adverse-effect level robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating no-observed-adverse-effect level robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computational Variant Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of computational variant evidence.
bemo
BEMO:2000215
BEMO:2000215
222
1f84faa3d8c87f3b1ed44b62077d2c489782ab72b149d6009c127395a0070a10
8
Assesses computational variant evidence using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in computational variant evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating computational variant evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Cross-Omics Replication
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cross-omics replication.
bemo
BEMO:2000305
BEMO:2000305
312
5b9be25d7af680eef194d1d2e393dbdeaaf607f62c516cb41c75f5dfd1ba50b9
10
Assesses cross-omics replication using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in cross-omics replication can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cross-omics replication as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Developing
Computation profile for Intervention Fidelity in Animal Studies
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Reference Standard Validity
The degree to which reference standard supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000134
BEMO:2000134
141
3bd12b9aed8e6e227ec8b1e5082fc9060b988babbc4204f8a360bb7503e6ad49
10
Assesses reference standard validity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in reference standard validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reference standard validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Internal Standard Performance
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of internal standard performance.
bemo
BEMO:2000365
BEMO:2000365
372
904dce77d19fcc637d6eae56d2d92b19b7052fd578de120e2a18bb08d00f2f54
8
Assesses internal standard performance using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in internal standard performance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating internal standard performance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Ontology Annotation Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 368
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Quantification Accuracy; Missing-Value Burden; Ion Suppression Assessment", "Common Misinterpretations": "Treating dynamic range coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Dynamic Range Coverage", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The proportion and representativeness of the relevant dynamic range captured by the evidence or measurement process.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses dynamic range coverage using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in dynamic range coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
368
57e4c77ea2a48565f0758aadf8529c7db52bab9beb67f6e9c8366f0f3232e753
Phenotypic Concordance
The degree of agreement in phenotypic across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000021
BEMO:2000021
28
d6b2622e1418a4a4c75dd4e74edee6d3fe7ea5ecac7eb6c453c614421e29173a
8
Assesses phenotypic concordance using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in phenotypic concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating phenotypic concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Case-Level Evidence Strength
The magnitude and credibility of independent evidence supporting case-level evidence.
bemo
BEMO:2000213
BEMO:2000213
220
e9e681d8a3c5b4ca82348a6cd8fe70f3f48bd8351d482c3aa492a6712873c3db
8
Assesses case-level evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in case-level evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating case-level evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Collider Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Time-Varying Confounding Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of time-varying confounding control.
bemo
BEMO:2000108
BEMO:2000108
115
f46690f0cea7611ea5bb0e9a0e5c1b90d25ecad9c8f28f3083cc2a71b53ce7f0
10
Assesses time-varying confounding control using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in time-varying confounding control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.
Treating time-varying confounding control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Network Context Support
The magnitude and credibility of independent evidence supporting network context.
bemo
BEMO:2000016
BEMO:2000016
23
d14cae7a1ae4e812a18da95395619423ade3837794dc32beea80cb11af0003f0
8
Assesses network context support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in network context support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating network context support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Evidence Quality
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Hotspot/Functional-Domain Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of hotspot/functional-domain evidence.
bemo
BEMO:2000223
BEMO:2000223
230
d2e72ca269dc077aef430f8efab1ad3e7adf8db5e6aa339468796f69dfaaee12
8
Assesses hotspot/functional-domain evidence using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in hotspot/functional-domain evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating hotspot/functional-domain evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 176
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Animal Model Construct Validity; Species Appropriateness; Sex as a Biological Variable Adequacy", "Common Misinterpretations": "Treating animal model predictive validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Animal Model Predictive Validity", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The degree to which animal model predictive supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses animal model predictive validity using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in animal model predictive validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
176
ae7f68af5b958c01b88066cf76664c77d333abaff4414b63760b817bf19359d0
Variant Phase Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of variant phase evidence.
bemo
BEMO:2000236
BEMO:2000236
243
e63350cbc225389693f86f1ceb6c24e47930b1a6b522574479bcaf1390376503
8
Assesses variant phase evidence using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in variant phase evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating variant phase evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Missing-Data Reporting Completeness
The extent to which all scientifically necessary components of missing-data reporting are present, documented, and evaluable.
bemo
BEMO:2000418
BEMO:2000418
425
810e02862dfc4630d6f7a0d7bd830ccd305d847e2d0153cfdb863b68f1c2ece8
9
Assesses missing-data reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in missing-data reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating missing-data reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Variant Call Quality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of variant call quality.
bemo
BEMO:2000266
BEMO:2000266
273
45320c5b96fc37b8148ee2721453c92e6bec79491a4ed2d2f6fedbc2cb99cb01
8
Assesses variant call quality using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in variant call quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating variant call quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Positive Likelihood Ratio
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of positive likelihood ratio.
bemo
BEMO:2000127
BEMO:2000127
134
65ab6bc3c7bfa89dbab942e4977888e1ec5d3ea554c360ceffa76f4daab1fd7e
8
Assesses positive likelihood ratio using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in positive likelihood ratio can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.
Treating positive likelihood ratio as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Sex as a Biological Variable Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Population Applicability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Mediation Evidence Strength
The magnitude and credibility of independent evidence supporting mediation evidence.
bemo
BEMO:2000099
BEMO:2000099
106
b7abd5f34cd2addb34f26d923a9c80b402d8914f57510a7f5fe8c4ed41659c58
8
Assesses mediation evidence strength using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in mediation evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating mediation evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Case-Control Evidence Strength
The magnitude and credibility of independent evidence supporting case-control evidence.
bemo
BEMO:2000212
BEMO:2000212
219
a2868c878aff5e68e6f9aeb87fd5a98fd776f1ca106d2e9dbad34a8cc81ecfd8
8
Assesses case-control evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in case-control evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating case-control evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Site-to-Site Assay Portability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of site-to-site assay portability.
bemo
BEMO:2000301
BEMO:2000301
308
c0e7317023f6c1fc290d021e4b38163503a72ee411cc671fe6aae25ee84d1ee5
8
Assesses site-to-site assay portability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in site-to-site assay portability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating site-to-site assay portability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Effect Magnitude
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of effect magnitude.
bemo
BEMO:2000439
BEMO:2000439
446
e21896fd55cc8da1ecb273cc15de195832d87a4d4554d2b68b3a9ce060442d84
8
Assesses effect magnitude using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in effect magnitude can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating effect magnitude as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Content Validity
The degree to which content supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000036
BEMO:2000036
43
1ceb580fd3a672a7c14d44163bda2886dd03d14b99ce004f1db40aa485bb3e1b
10
Assesses content validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in content validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating content validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Biomarker-Outcome Association Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Sequencing Depth Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Specification-Curve Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of specification-curve robustness.
bemo
BEMO:2000406
BEMO:2000406
413
290b2551a20c17c23c05ae84270ce7ac2c9d4a43b9817050cdbe8a353ce743ac
8
Assesses specification-curve robustness using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in specification-curve robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating specification-curve robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Specialized / infrequent
Developing
Computation profile for Transport Condition Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Demographic Generalizability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of demographic generalizability.
bemo
BEMO:2000194
BEMO:2000194
201
e22341c32d6f7aad19740ed4b08e7f5ffc168107f1e04a79ba4ec60f6b92c463
8
Assesses demographic generalizability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in demographic generalizability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating demographic generalizability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Negative Likelihood Ratio
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of negative likelihood ratio.
bemo
BEMO:2000122
BEMO:2000122
129
6dc29889e79017542636c8b841aff4c0f5947b19dec51307e00a9c51cdc593cc
8
Assesses negative likelihood ratio using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in negative likelihood ratio can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.
Treating negative likelihood ratio as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Code-Sharing Transparency
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of code-sharing transparency.
bemo
BEMO:2000407
BEMO:2000407
414
d85d4b116984de9e36c9410fc248aae4b536153c50f4651cf7e48c6b558b49eb
8
Assesses code-sharing transparency using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in code-sharing transparency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating code-sharing transparency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Reproducibility Information Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Partial Area Under the Receiver Operating Characteristic Curve
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Discrimination of Statistical Predictions
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of discrimination of statistical predictions.
bemo
BEMO:2000437
BEMO:2000437
444
cf100cd63561a5d7a167b9d561a196b9309ca048456a8705766e76a54a524fce
8
Assesses discrimination of statistical predictions using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in discrimination of statistical predictions can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating discrimination of statistical predictions as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pharmacokinetic Adequacy
The extent to which pharmacokinetic is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000347
BEMO:2000347
354
2b02efdd0f6a774096cee910d90b575b8f28f136a8ee594876642fb213f66330
8
Assesses pharmacokinetic adequacy using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in pharmacokinetic adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating pharmacokinetic adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Multiverse Analysis Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of multiverse analysis robustness.
bemo
BEMO:2000400
BEMO:2000400
407
80472d064c943a6566bee4bd605245ecac370ebde0eb5ae577d7687482d47cc8
8
Assesses multiverse analysis robustness using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in multiverse analysis robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating multiverse analysis robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Specialized / infrequent
Developing
Computation profile for Metabolite Annotation Level
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Effect-Modification Credibility
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of effect-modification credibility.
bemo
BEMO:2000094
BEMO:2000094
101
f66521085db07e7bd17f7b93b07cf74e57b1c85461a2555c1146f672069098b2
8
Assesses effect-modification credibility using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in effect-modification credibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating effect-modification credibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Tumor Purity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Matrix Effect
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of matrix effect.
bemo
BEMO:2000288
BEMO:2000288
295
e901d4ca31939164c2ecbe0d0bed735b156c8e09a3c3f28444f5359947798b49
8
Assesses matrix effect using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in matrix effect can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating matrix effect as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 346
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Species Extrapolation Validity; Mixture Interaction Assessment", "Common Misinterpretations": "Treating human-relevance of toxicological evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Human-Relevance of Toxicological Evidence", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of human-relevance of toxicological evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses human-relevance of toxicological evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in human-relevance of toxicological evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
346
32b4aadfeeab004a7a018b1fe09f04f0b2283eb92e7749123f8ac828db8cfce6
Computation profile for Conflicting Interpretation Burden
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Off-Target Liability Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of off-target liability evidence.
bemo
BEMO:2000017
BEMO:2000017
24
688a2bdbd29ad4b59eca6350cf3462f2bc158ee3d70ef4de59f338b6e6c797a6
8
Assesses off-target liability evidence using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in off-target liability evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating off-target liability evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Missing-Data Mechanism Plausibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Genotype Quality
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 249
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Library Complexity; Sample Identity Concordance; Sex Concordance", "Common Misinterpretations": "Treating contamination burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Contamination Burden", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of contamination burden.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses contamination burden using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in contamination burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
249
fc0cd98bc92235d469bf2bc025f0d327d8170dfb4df97b973710a8d887eff3f4
Computation profile for Human-Relevance of Toxicological Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Type I Error Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Variant Pathogenicity Evidence Strength
The magnitude and credibility of independent evidence supporting variant pathogenicity evidence.
bemo
BEMO:2000235
BEMO:2000235
242
ce239eedae0622eb1018ab2bfc52cd2d32ef0a5f6a874008c2d8838deb99bd4a
8
Assesses variant pathogenicity evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in variant pathogenicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating variant pathogenicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Quantification Precision
The closeness of repeated estimates or measurements and the narrowness of uncertainty around quantification.
bemo
BEMO:2000382
BEMO:2000382
389
6e5af7c1fffa7603da0abd7518a670fd09cade5fe3e47c8417886f91e6b4e07a
9
Assesses quantification precision using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in quantification precision can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating quantification precision as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Network Node Confidence
The justified degree of certainty assigned to network node given the quantity, quality, consistency, and limitations of supporting evidence.
bemo
BEMO:2000316
BEMO:2000316
323
79eaf0a8e6c6fc3012d48e379c545c17d89aecf6de033d40ae6399ec89441a3b
10
Assesses network node confidence using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in network node confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating network node confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Proteome Coverage
The proportion and representativeness of the relevant proteome captured by the evidence or measurement process.
bemo
BEMO:2000380
BEMO:2000380
387
5fec1390446b523e8b6072b1babcb11902375d326650ad448b800e258e0a8401
9
Assesses proteome coverage using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in proteome coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating proteome coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Harms Reporting Completeness
The extent to which all scientifically necessary components of harms reporting are present, documented, and evaluable.
bemo
BEMO:2000414
BEMO:2000414
421
15ffc51fd66113f5e5c480a7127c592104dbff4a2e6faafcfafb509cd725ea07
9
Assesses harms reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in harms reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating harms reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 474
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Baseline Comparability; Performance Bias Risk; Detection Bias Risk", "Common Misinterpretations": "Treating blinding integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Blinding Integrity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of blinding integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses blinding integrity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in blinding integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
474
4e84b980ded4e833e93013ba317d5857e2211ad9b612e4b77eb9308cf8244573
Computation profile for Population Representativeness
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Causal Contrast Clarity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of causal contrast clarity.
bemo
BEMO:2000085
BEMO:2000085
92
d81e7f98cb339c656ef93d3243144295a2c872dc2bcb50f64863414984f27993
10
Assesses causal contrast clarity using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in causal contrast clarity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.
Treating causal contrast clarity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Metadata Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Sampling Frame Adequacy
The extent to which sampling frame is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000205
BEMO:2000205
212
39b818e971de479787fde11a4523a1b0d31f5dcbc88f645b968cff60fa6bf381
8
Assesses sampling frame adequacy using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in sampling frame adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating sampling frame adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Target Engagement Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Study Design Appropriateness
The extent to which study design is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000487
BEMO:2000487
494
7f9fb8e0c1ccb555a18bd7b6b07b3173d661a6707fab17a21ee632e59503b393
8
Assesses study design appropriateness using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in study design appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating study design appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Exclusion-Criteria Prespecification
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of exclusion-criteria prespecification.
bemo
BEMO:2000174
BEMO:2000174
181
9718ea523adf00e3cd776b8e22423abc28cf080bd659786c7568ddd524ab9208
8
Assesses exclusion-criteria prespecification using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in exclusion-criteria prespecification can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating exclusion-criteria prespecification as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Disease-Severity Generalizability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of disease-severity generalizability.
bemo
BEMO:2000195
BEMO:2000195
202
867ff07055e7596b98039def11c9ce9ca056af7a3a11134bb8f46a0bd5adc97a
8
Assesses disease-severity generalizability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in disease-severity generalizability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating disease-severity generalizability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 289
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Reagent Lot Consistency; Operator Variability; Site-to-Site Assay Portability", "Common Misinterpretations": "Treating instrument drift as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Instrument Drift", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of instrument drift.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses instrument drift using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in instrument drift can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
289
39eec82dab659b7826d3bea8f3227e8afeb55e0f15214b02955200cb8a59a8ba
Computation profile for Freeze–Thaw Burden
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Evidence Coherence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence coherence.
bemo
BEMO:2000142
BEMO:2000142
149
573628981de9da600edff4f13001018973ec4403d9ad58179a02f6958d0404dd
8
Assesses evidence coherence using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence coherence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence coherence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Receptor Occupancy Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of receptor occupancy evidence.
bemo
BEMO:2000350
BEMO:2000350
357
5025e34702ce03c8ed7e7aa3ed95dd9bde6bac1786183651ce608b34ffcc6a28
8
Assesses receptor occupancy evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in receptor occupancy evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating receptor occupancy evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Spatial Biological Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Toxicokinetic Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Conflict-of-Interest Transparency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Cumulative Evidence Stability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
DNA Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of dna integrity.
bemo
BEMO:2000065
BEMO:2000065
72
494036e6478a34bfe608bc3b1267e04a777d16e96430109ab2ba9b01e99c6693
8
Assesses dna integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in dna integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating dna integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 365
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Benchmark Dose Reliability; Adverse Outcome Pathway Support; Organ-Specific Toxicity Evidence", "Common Misinterpretations": "Treating toxicological mode-of-action support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Toxicological Mode-of-Action Support", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The magnitude and credibility of independent evidence supporting toxicological mode-of-action.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses toxicological mode-of-action support using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in toxicological mode-of-action support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
365
05b74bda1c6ce0d8b99d4d0ffdc239f28a207c0631f25d1605f8d1438c74148a
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 393
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Conceptual Replication Success; Computational Reproducibility; Experimental Reproducibility", "Common Misinterpretations": "Treating analytical reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Analytical Reproducibility", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The degree to which analytical reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses analytical reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in analytical reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
393
8e400bbae5ea25bca3af45f9c956cb2ceeb52ded61247de3b92d5364736bab99
Network Reconstruction Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of network reconstruction robustness.
bemo
BEMO:2000317
BEMO:2000317
324
0b4f2bdfa3ccc3b3abdb67b1f3be15b6d8479743942152580ad3f1929db67d5f
8
Assesses network reconstruction robustness using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in network reconstruction robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating network reconstruction robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 464
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Posterior Probability Strength; Equivalence Margin Validity; Noninferiority Margin Validity", "Common Misinterpretations": "Treating prior sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Prior Sensitivity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prior sensitivity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses prior sensitivity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in prior sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
464
ced026349740f092de03ee4146894b4367dab3b43793f3aa048eac1fd782cbda
Pharmacodynamic Biomarker Validity
The degree to which pharmacodynamic biomarker supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000050
BEMO:2000050
57
4eb497c772907b7927e32ee929fba9a804b1c75888594041caa61dfab977440a
10
Assesses pharmacodynamic biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in pharmacodynamic biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating pharmacodynamic biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Differential Verification Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Reproductive Toxicity Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Cumulative Evidence Stability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cumulative evidence stability.
bemo
BEMO:2000141
BEMO:2000141
148
e87943e84aae44a1eeb6300f5a52416f0528fe0f3bf76f40f0a2a4486c87ea6a
9
Assesses cumulative evidence stability using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in cumulative evidence stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cumulative evidence stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Functional Variant Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Pharmacodynamic Biomarker Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Evidence Precision
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Accuracy
The closeness of accuracy to the accepted reference or true value.
bemo
BEMO:2000267
BEMO:2000267
274
35d988975cf78a06cb70788f89a4196551552de36a3cf47e57d2d96fc8cf41fd
8
Assesses accuracy using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Pharmacokinetic Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Comparator Description Completeness
The extent to which all scientifically necessary components of comparator description are present, documented, and evaluable.
bemo
BEMO:2000408
BEMO:2000408
415
a4617b71362b91773383e09e14a13eab2a1f342da234aa0f9bf7cc4400a6ff10
9
Assesses comparator description completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in comparator description completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating comparator description completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Population Applicability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of population applicability.
bemo
BEMO:2000202
BEMO:2000202
209
9073b7f8b73fc788aa084d5c4ac174d69482c2a1439b22f86a39e465503328bc
8
Assesses population applicability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in population applicability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating population applicability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 177
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Humane Endpoint Appropriateness; Exclusion-Criteria Prespecification; Experimental Batch Randomization", "Common Misinterpretations": "Treating attrition accounting in animal studies as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Attrition Accounting in Animal Studies", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of attrition accounting in animal studies.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses attrition accounting in animal studies using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in attrition accounting in animal studies can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
177
bfb6735bf1e0494f94802ec087edc274ccdaea50558f3ea043f6f21ae9210505
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 141
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Threshold Validity; Index-Test Blinding; Verification Bias Risk", "Common Misinterpretations": "Treating reference standard validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Reference Standard Validity", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "The degree to which reference standard supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses reference standard validity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reference standard validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
141
3bd12b9aed8e6e227ec8b1e5082fc9060b988babbc4204f8a360bb7503e6ad49
Computation profile for No-Interference Plausibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Analytical Measurement Range
0.1.0
Apply a validated analyte- and method-specific measurement procedure with calibration and quality control.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","unitRef":"REQUIRED"}
specimen_or_material; measurement_procedure; calibration_reference; quality_control_results; unit
replicate_measurements; environmental_conditions; instrument_version
xsd:decimal
Method- and analyte-specific physical units
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required.
Source maturity: Established; BEMO computation profile requires independent validation.
Ontology Evidence-Code Strength
The magnitude and credibility of independent evidence supporting ontology evidence-code.
bemo
BEMO:2000319
BEMO:2000319
326
98cc3b5d977b0d0bd7a309c477d2653f05eeb7554078021a2d0b2bf21036d891
8
Assesses ontology evidence-code strength using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in ontology evidence-code strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating ontology evidence-code strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Biomarker Qualification Strength
The magnitude and credibility of independent evidence supporting biomarker qualification.
bemo
BEMO:2000029
BEMO:2000029
36
35003a5d4be2844c7914597bb95d8d44ce388ce934a6b5067c8e498ac9e8de35
8
Assesses biomarker qualification strength using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in biomarker qualification strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biomarker qualification strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Co-intervention Bias Risk
The probability or degree that co-intervention bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000470
BEMO:2000470
477
de36d6b209ddc70bf44fd968006b0ba61482bcc4bbe42db3c2a605301631de5e
10
Assesses co-intervention bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in co-intervention bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating co-intervention bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Network Edge Confidence
The justified degree of certainty assigned to network edge given the quantity, quality, consistency, and limitations of supporting evidence.
bemo
BEMO:2000315
BEMO:2000315
322
6118b228a2af4e149e82150fb026ef08cdd6d05b9748a0128d6e8c14f147784c
10
Assesses network edge confidence using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in network edge confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating network edge confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Proteoform Identification Confidence
The justified degree of certainty assigned to proteoform identification given the quantity, quality, consistency, and limitations of supporting evidence.
bemo
BEMO:2000379
BEMO:2000379
386
8df86e86625f432ba76b927808300a9cfa4a9b6844cae40ca0c6240f4f899eda
10
Assesses proteoform identification confidence using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in proteoform identification confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating proteoform identification confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Predictive Biomarker Validity
The degree to which predictive biomarker supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000051
BEMO:2000051
58
c5c7cbb5b7738940f61731df207186f05f8ae71ecab7ca3c5588c823af8d00eb
10
Assesses predictive biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in predictive biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating predictive biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Criterion Validity
The degree to which criterion supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000039
BEMO:2000039
46
04af58cc2d88f851cf92e0d1960d29a74b3d8a5cf3525dda9579101a813732b2
10
Assesses criterion validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in criterion validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating criterion validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Pharmacodynamic Adequacy
The extent to which pharmacodynamic is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000346
BEMO:2000346
353
99eef93672d2993b0ec131c09788c14463fd2a5166870e81b6b8d81c3637506b
8
Assesses pharmacodynamic adequacy using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in pharmacodynamic adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating pharmacodynamic adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 36
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Context-of-Use Validity; Biomarker Reliability; Biomarker Responsiveness", "Common Misinterpretations": "Treating biomarker qualification strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Biomarker Qualification Strength", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The magnitude and credibility of independent evidence supporting biomarker qualification.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses biomarker qualification strength using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biomarker qualification strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
36
35003a5d4be2844c7914597bb95d8d44ce388ce934a6b5067c8e498ac9e8de35
Target Trial Emulation Fidelity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of target trial emulation fidelity.
bemo
BEMO:2000107
BEMO:2000107
114
7a907a8dc4fe2c40df0aa3c333c85c8e7231c1e33d26321171335bcd3464b6a6
8
Assesses target trial emulation fidelity using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in target trial emulation fidelity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating target trial emulation fidelity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Analytical Reproducibility
The degree to which analytical reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000386
BEMO:2000386
393
8e400bbae5ea25bca3af45f9c956cb2ceeb52ded61247de3b92d5364736bab99
10
Assesses analytical reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in analytical reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating analytical reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Variant Phase Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Analytical Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Knowledge-Graph Provenance Quality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of knowledge-graph provenance quality.
bemo
BEMO:2000311
BEMO:2000311
318
50aaf6749314a5ad4694891f48d60c74feaa99da022b5cc91c17f4221f0696a2
8
Assesses knowledge-graph provenance quality using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in knowledge-graph provenance quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating knowledge-graph provenance quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Developing
Model Specification Adequacy
The extent to which model specification is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000450
BEMO:2000450
457
c5d48404d8de034fc4aa45bb9f64638b1d861e61321423aa5f11d4f63cc6783b
8
Assesses model specification adequacy using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in model specification adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating model specification adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Consistency Assumption Plausibility
The degree of agreement in consistency assumption plausibility across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000089
BEMO:2000089
96
1bec3d5308662b9b5fad28978574efe7573d32766cf45ac34286863d2d2267a8
9
Assesses consistency assumption plausibility using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in consistency assumption plausibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating consistency assumption plausibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Variant Classification Stability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Raw-Data Availability
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Biological Plausibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Reverse-Causation Risk
The probability or degree that reverse-causation introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000105
BEMO:2000105
112
17fc4b74f3434a4ae22f908b09530b9a1be33f82af499c89239235d78d7e1375
8
Assesses reverse-causation risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in reverse-causation risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reverse-causation risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Model Fit
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Loss-of-Function Validation
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of loss-of-function validation.
bemo
BEMO:2000008
BEMO:2000008
15
7b09e166eefa32f832c89264de2b3521f93b540fbc674d78fc625fd792779d09
8
Assesses loss-of-function validation using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in loss-of-function validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating loss-of-function validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Expressivity Consistency
The degree of agreement in expressivity across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000218
BEMO:2000218
225
d100842d4d2c367ecd1122df00784975aa0a773a92bbb5c68a8ff23fbbd4dfc4
8
Assesses expressivity consistency using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in expressivity consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating expressivity consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Splicing Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Missing Evidence Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Biomarker and Endpoint Validation metric
Category of biomedical evidence metrics concerned with biomarker and endpoint validation.
bemo
BEMO:1100002
Candidate
Computation profile for Disease-Severity Generalizability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Biomarker Responsiveness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biomarker responsiveness.
bemo
BEMO:2000031
BEMO:2000031
38
19a54854b5dfb7168eddad304da95ddbab2e1ad34992146fcedfd650dfe09548
8
Assesses biomarker responsiveness using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in biomarker responsiveness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biomarker responsiveness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Biospecimen Provenance Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 367
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Extraction Recovery; Metabolic Feature Reproducibility; Pathway Enrichment Robustness", "Common Misinterpretations": "Treating derivatization efficiency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Derivatization Efficiency", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of derivatization efficiency.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses derivatization efficiency using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in derivatization efficiency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
367
d83b78fc8beda508a5c18da8dbb9e284c948a9741d4fb8d070f9f30bc9267827
Mass Accuracy
The closeness of mass to the accepted reference or true value.
bemo
BEMO:2000368
BEMO:2000368
375
ebb0e828e6a056d20bce6788aee10b743cfcc2fc175a82fd693585711880ff73
8
Assesses mass accuracy using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in mass accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating mass accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Processed-Data Availability
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 456
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Model Specification Adequacy; Residual Diagnostics Adequacy; Distributional Assumption Adequacy", "Common Misinterpretations": "Treating model fit as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Model Fit", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of model fit.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses model fit using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in model fit can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
456
bf6f70bdb1883a9c1ba70dc7d37eb1885dd46ed70094a2b12bd2cc41e0f53bad
Metabolic Feature Reproducibility
The degree to which metabolic feature reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000369
BEMO:2000369
376
521bf8b7c1bbe892685e95c05cc7d63a05c4e26119e9951d705ff460b3bb99d1
10
Assesses metabolic feature reproducibility using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in metabolic feature reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating metabolic feature reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Negative Likelihood Ratio
0.1.0
value = numerator / denominator; the null value is typically 1 where scientifically applicable
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator != 0"],"null_value":1}
numerator; denominator; operational_definition; assessment_context
confidence_level; stratum
xsd:decimal
Ratio scale; null typically 1
0.0
1
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Not generally required unless converted to a probability or score.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Reportable Range
0.1.0
Apply a validated analyte- and method-specific measurement procedure with calibration and quality control.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","unitRef":"REQUIRED"}
specimen_or_material; measurement_procedure; calibration_reference; quality_control_results; unit
replicate_measurements; environmental_conditions; instrument_version
xsd:decimal
Method- and analyte-specific physical units
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required.
Source maturity: Established; BEMO computation profile requires independent validation.
Publication Bias Risk
The probability or degree that publication bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000162
BEMO:2000162
169
ae14a6a3126f2ecd31144006a017263d5e10df0fcde0e6bd8cc0872e93853a7a
10
Assesses publication bias risk using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in publication bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Forest plots; heterogeneity statistics; tau-squared; prediction intervals; funnel plots; regression or selection models; sensitivity analyses.
Treating publication bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Follow-up Completeness
The extent to which all scientifically necessary components of follow-up are present, documented, and evaluable.
bemo
BEMO:2000478
BEMO:2000478
485
a99fe581309ce29c4e888ffbd9a1cba6c99a6948f2a8e99e23714e0f40920b87
9
Assesses follow-up completeness using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in follow-up completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating follow-up completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Specification-Curve Robustness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Developing; BEMO computation profile requires independent validation.
Temporal Biological Concordance
The degree of agreement in temporal biological across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000026
BEMO:2000026
33
57d68c3c49543e85a4a6ee1a26a1705e9e0e05df169673d5a8f0a00632bfa39b
8
Assesses temporal biological concordance using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in temporal biological concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating temporal biological concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Perturbational Validation
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
External Validity
The degree to which external supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000197
BEMO:2000197
204
291600cbcf719af0c4e82629e00cccabf1b55b19eff672c1fb6c5c7fe2ec4338
10
Assesses external validity using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in external validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating external validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Extraction Recovery
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 247
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Reference Bias; Hardy–Weinberg Equilibrium Compatibility; Batch-Effect Control", "Common Misinterpretations": "Treating call-rate completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Call-Rate Completeness", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "The extent to which all scientifically necessary components of call-rate are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses call-rate completeness using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in call-rate completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
247
c69a25da4615faa5675944ca32d4c5ec30a0c44f5216034581329797d9a32abc
Dynamical Stability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of dynamical stability.
bemo
BEMO:2000306
BEMO:2000306
313
d58852cf510c70b8fcfa12f792791c36728adc16d59c5b9bb4b7a9b129c39f52
9
Assesses dynamical stability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in dynamical stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating dynamical stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Gene–Disease Validity
The degree to which gene–disease supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000221
BEMO:2000221
228
4501a0d637c0d3d7b3852d3be8f61115b2b1ac442582dd9a3ec88bd21dad9f7c
10
Assesses gene–disease validity using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in gene–disease validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating gene–disease validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Intervention Fidelity in Animal Studies
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of intervention fidelity in animal studies.
bemo
BEMO:2000180
BEMO:2000180
187
f4ebc6af400b7084f27d64f43d25beed78acdab6caf4a03c3179cab231054034
8
Assesses intervention fidelity in animal studies using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in intervention fidelity in animal studies can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating intervention fidelity in animal studies as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Study Design, Statistical Validity, and Causal Inference metric
Metrics assessing internal validity, statistical inference, causal identification, bias, and study design.
bemo
BEMO:1000002
Candidate
Computation profile for Between-Study Variance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Genetic Background Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of genetic background control.
bemo
BEMO:2000177
BEMO:2000177
184
3f6791503c2895971802b5fb5573311999fc32cde51839dc771c56e4855ceb83
8
Assesses genetic background control using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in genetic background control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating genetic background control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Confidence Interval Compatibility
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Alternative Splicing Validation
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Conflict-of-Interest Transparency
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of conflict-of-interest transparency.
bemo
BEMO:2000409
BEMO:2000409
416
e951e1891736a22e6ef317b85f2f1f3032c69473b10d7ac72b1ed8342656df25
8
Assesses conflict-of-interest transparency using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in conflict-of-interest transparency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating conflict-of-interest transparency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 175
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Blinding in Experimental Assessment; Animal Model Construct Validity; Animal Model Predictive Validity", "Common Misinterpretations": "Treating animal model face validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Animal Model Face Validity", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The degree to which animal model face supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses animal model face validity using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in animal model face validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
175
1bc9d94635678f728d2a96fd7016f3d61e24c78d9e33c5093213c9d715511534
Computation profile for RNA Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Type II Error Risk
The probability or degree that type ii error introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000461
BEMO:2000461
468
de4287424d92a6f5a119531d8abefacf3d79beff40059540425ff1afc5f1021d
8
Assesses type ii error risk using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in type ii error risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating type ii error risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Negative-Control Performance
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of negative-control performance.
bemo
BEMO:2000181
BEMO:2000181
188
0095c2122f8f2e2a2aa6070b8c33f6576321db691574400350300c93ce956830
8
Assesses negative-control performance using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in negative-control performance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating negative-control performance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Spectrum Representativeness
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 450
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Missing-Data Sensitivity; Outlier Influence Robustness; Influential Observation Sensitivity", "Common Misinterpretations": "Treating imputation validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Imputation Validity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The degree to which imputation supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses imputation validity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in imputation validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
450
451bb1f13aa8cd7feb728a231f72c948355ace1a73de7f1903381352517ceaf5
Data Provenance Completeness
The extent to which all scientifically necessary components of data provenance are present, documented, and evaluable.
bemo
BEMO:2000391
BEMO:2000391
398
da25c768c2ee710f0cf6b50b686a1880c8166dd651b4f062c14e952fc1bdf622
9
Assesses data provenance completeness using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in data provenance completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating data provenance completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Outcome Ascertainment Validity
The degree to which outcome ascertainment supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000480
BEMO:2000480
487
654f73aad6650cd58d63f5a49ae9e034bfcb22952cbf7a018f652c070e1ba6c1
10
Assesses outcome ascertainment validity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in outcome ascertainment validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating outcome ascertainment validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Cellularity Adequacy
The extent to which cellularity is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000061
BEMO:2000061
68
de5efa73121b798bb0e1a53bddae9ef743d4a522e6df52007a5b7569bc63e020
8
Assesses cellularity adequacy using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in cellularity adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cellularity adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Intervention Classification Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Reclassification Improvement
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of reclassification improvement.
bemo
BEMO:2000133
BEMO:2000133
140
74d83029c18dff3247c8fb9baae1d6aba52c9b7328d218dd0210458cbd5ad7d5
8
Assesses reclassification improvement using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in reclassification improvement can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reclassification improvement as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Study Design Appropriateness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Cold Ischemia Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Measurement Error Correction
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Population Representativeness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of population representativeness.
bemo
BEMO:2000203
BEMO:2000203
210
efa730876069251ed2dde336a59f381b8c109c4ee429da0b68bf8bae61b0ba48
8
Assesses population representativeness using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in population representativeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating population representativeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Content Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Observed-to-Expected Ratio
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Intervention Description Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 258
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Differential Expression Robustness; Transcript Quantification Reliability; Alternative Splicing Validation", "Common Misinterpretations": "Treating normalization adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Normalization Adequacy", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "The extent to which normalization is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses normalization adequacy using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in normalization adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
258
518f55960ed771c22a0c382ff436c097ae34f4481b35b5a64e7f2cf8493778a4
Biomarker-Outcome Association Strength
The magnitude and credibility of independent evidence supporting biomarker-outcome association.
bemo
BEMO:2000033
BEMO:2000033
40
d54499df144e3534026302c035690746b731c0abf25107f5fccdd8bada0bd79a
8
Assesses biomarker-outcome association strength using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in biomarker-outcome association strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biomarker-outcome association strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
metric category
A source-preserving organizational class grouping biomedical evidence metrics within a metric pillar.
bemo
BEMO:0000003
Candidate
Reagent Lot Consistency
The degree of agreement in reagent lot across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000294
BEMO:2000294
301
66143063e8139ffa9de4cc6fb03a0eb1778bd92be05c901387fc7af071b0e433
8
Assesses reagent lot consistency using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in reagent lot consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reagent lot consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Cross-Reactivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cross-reactivity.
bemo
BEMO:2000277
BEMO:2000277
284
1de0ffae2d83059ba1d111ea3bf309f20078d963ac7b7f912a42e8b4e10e5e32
8
Assesses cross-reactivity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in cross-reactivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cross-reactivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Fragmentation Spectrum Quality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of fragmentation spectrum quality.
bemo
BEMO:2000364
BEMO:2000364
371
4295f539b87882bc52461f7bd190419d6afdf7864927b9a24efe1e5f65c9fcd7
8
Assesses fragmentation spectrum quality using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in fragmentation spectrum quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating fragmentation spectrum quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Sample Size Justification
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of sample size justification.
bemo
BEMO:2000185
BEMO:2000185
192
9c4289123cb7c0f9d87fa9453151e360a9b6781cd9fbb6f6b26a09f4543b0dab
8
Assesses sample size justification using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in sample size justification can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating sample size justification as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
metric pillar
A high-level organizational class grouping biomedical evidence metrics by scientific function.
bemo
BEMO:0000002
Candidate
Computation profile for Batch-Effect Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Participant Flow Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Processing Delay Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Lowest-Observed-Adverse-Effect Level Robustness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Distributional Assumption Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Follow-up Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Internal Standard Performance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Warm Ischemia Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Verification Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Evidence Completeness
The extent to which all scientifically necessary components of evidence are present, documented, and evaluable.
bemo
BEMO:2000143
BEMO:2000143
150
263bd485141b0cc75ee285fb1989c667b83eb378666886a960ed5bd2a610f232
9
Assesses evidence completeness using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Prior Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prior sensitivity.
bemo
BEMO:2000457
BEMO:2000457
464
ced026349740f092de03ee4146894b4367dab3b43793f3aa048eac1fd782cbda
9
Assesses prior sensitivity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in prior sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating prior sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Chain-of-Custody Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of chain-of-custody integrity.
bemo
BEMO:2000062
BEMO:2000062
69
38ecb99023f4b12b1bfca7d66221917ae384546ef170c868cd32fb2eb70f3f40
8
Assesses chain-of-custody integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in chain-of-custody integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating chain-of-custody integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Perturbational Validation
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of perturbational validation.
bemo
BEMO:2000020
BEMO:2000020
27
e56c259480c9a3f5c7efe467a62a9721824469292b9aeb900d1804f2bbfb7983
8
Assesses perturbational validation using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in perturbational validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating perturbational validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 157
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Directness; Evidence Coherence; Evidence Consensus Strength", "Common Misinterpretations": "Treating evidence precision as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Precision", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The closeness of repeated estimates or measurements and the narrowness of uncertainty around evidence.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses evidence precision using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence precision can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
157
fed12d9ff25f08fc6db399dc5bed2e8d719fc0db8bcefb33987a3ef3db40f55c
Intervention Classification Validity
The degree to which intervention classification supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000479
BEMO:2000479
486
283a3544a9dd5cb1a16190c929c82c6db2608db0f4c24860780498cb89bc0ac5
10
Assesses intervention classification validity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in intervention classification validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating intervention classification validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Cell-Type Specificity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cell-type specificity.
bemo
BEMO:2000003
BEMO:2000003
10
8fcb10135485d0dff26bef51519f2178a09a543c766f354a728ca79b2c39c582
9
Assesses cell-type specificity using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in cell-type specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cell-type specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Spatial Transcriptomic Registration Accuracy
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Assumption Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Detection Bias Risk
The probability or degree that detection bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000474
BEMO:2000474
481
98c36b49fe475eec66359fae3b2440fc1429c378e2bbe5896e082ec203855535
10
Assesses detection bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in detection bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating detection bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Computation profile for Collection Procedure Consistency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Adverse Outcome Pathway Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Diagnostic Specificity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of diagnostic specificity.
bemo
BEMO:2000117
BEMO:2000117
124
14be8f0b88f11902e6b1562a38b9cb710840858fc59f748048f96b74c8f6c2b2
9
Assesses diagnostic specificity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in diagnostic specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.
Treating diagnostic specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Cross-Omics Replication
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Developing; BEMO computation profile requires independent validation.
Computation profile for Result Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Information Size Adequacy
The extent to which information size is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000157
BEMO:2000157
164
1349313eaf86ad90454f8dbf26aa9cb5389132fc9ac2e44b703024851df5136c
8
Assesses information size adequacy using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in information size adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating information size adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Measurement Uncertainty
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of measurement uncertainty.
bemo
BEMO:2000289
BEMO:2000289
296
02b8d5d020f47fcd1190b0fa52a15a86a5dd04748bbf2908e8c4e87ccdb4670f
10
Assesses measurement uncertainty using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in measurement uncertainty can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating measurement uncertainty as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Accuracy
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 92
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Correct Temporal Ordering; Target Trial Emulation Fidelity; Instrument Validity", "Common Misinterpretations": "Treating causal contrast clarity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Causal Contrast Clarity", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of causal contrast clarity.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses causal contrast clarity using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in causal contrast clarity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
92
d81e7f98cb339c656ef93d3243144295a2c872dc2bcb50f64863414984f27993
Computation profile for Predictive Biomarker Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Developmental Toxicity Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Evidence Consensus Strength
The magnitude and credibility of independent evidence supporting evidence consensus.
bemo
BEMO:2000145
BEMO:2000145
152
f207ed4da0508b76cb7e86f43fc62efc4861d327760c860608fc73ff3f4a6da8
8
Assesses evidence consensus strength using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence consensus strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence consensus strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Population Stratification Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of population stratification control.
bemo
BEMO:2000252
BEMO:2000252
259
c82d16c5b00c0127f5b02c510a48d677f57d7580cba4dcb68dce1141e1686c1d
8
Assesses population stratification control using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in population stratification control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating population stratification control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Period Effect Risk
The probability or degree that period effect introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000482
BEMO:2000482
489
78e4475c3f570e4a3e966d2bc7b940517b5d7eaa514d888dc22a6ea5b68d38fb
8
Assesses period effect risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in period effect risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating period effect risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Lowest-Observed-Adverse-Effect Level Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of lowest-observed-adverse-effect level robustness.
bemo
BEMO:2000341
BEMO:2000341
348
cd00c7765086fbbe65040e649de98f22bdf10298c91d8da0f6e7300a18f0ba9f
8
Assesses lowest-observed-adverse-effect level robustness using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in lowest-observed-adverse-effect level robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating lowest-observed-adverse-effect level robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Biological Replicate Adequacy
The extent to which biological replicate is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000171
BEMO:2000171
178
74976e4cf57f53c444f4721c04fc26fb7298cad4ac743db45feeb9a0be34587a
8
Assesses biological replicate adequacy using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in biological replicate adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biological replicate adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Immortal-Time Bias Risk
The probability or degree that immortal-time bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000097
BEMO:2000097
104
21a510319cb99a940f9fc99111522f80932b8a874f1ea984fb332f8f218483c8
10
Assesses immortal-time bias risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in immortal-time bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating immortal-time bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Counterevidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Evidence Triangulation Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Interlaboratory Reproducibility
The degree to which interlaboratory reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000396
BEMO:2000396
403
0977ae8d7278730723c8d024585a95d1125aa9fe6da897fe977a079b172ddb48
10
Assesses interlaboratory reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in interlaboratory reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating interlaboratory reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Co-segregation Likelihood
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Prognostic Biomarker Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 277
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Reproducibility of Measurement; Analytical Specificity; Limit of Detection", "Common Misinterpretations": "Treating analytical sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Analytical Sensitivity", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of analytical sensitivity.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses analytical sensitivity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in analytical sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
277
193a9fb566dfa18750368d0657f7d1e8c3de7251eea014f0d99b1b3e73df3afa
Necrosis Burden
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of necrosis burden.
bemo
BEMO:2000073
BEMO:2000073
80
33466881e023e8dc75d5f189be9da266f95f3bf8283fa46f9f4cff10e5af6f95
8
Assesses necrosis burden using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in necrosis burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating necrosis burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Observed-to-Expected Ratio
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of observed-to-expected ratio.
bemo
BEMO:2000124
BEMO:2000124
131
cdf33871a90f7426da799fbf8148672bff9e925c7981ff58372c39511bd4050e
8
Assesses observed-to-expected ratio using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in observed-to-expected ratio can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating observed-to-expected ratio as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Influential Observation Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Multiverse Analysis Robustness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Developing; BEMO computation profile requires independent validation.
Computation profile for De Novo Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Tissue Specificity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of tissue specificity.
bemo
BEMO:2000027
BEMO:2000027
34
e3cc6b49f3bcbfabd500e6222a7c3cc053aca55c824a6afddd04a9b448378651
9
Assesses tissue specificity using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in tissue specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating tissue specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Sex Concordance
The degree of agreement in sex across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000258
BEMO:2000258
265
897b2797d8e6449a88906be71f955c7726607e3b28833aef8661c12a890476a4
8
Assesses sex concordance using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in sex concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating sex concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Homeostatic Compensation Assessment
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of homeostatic compensation assessment.
bemo
BEMO:2000007
BEMO:2000007
14
61aaa3d7ba71760f01fcbbb79304031b61f416940944ef459f48b6d07dc0628d
8
Assesses homeostatic compensation assessment using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in homeostatic compensation assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating homeostatic compensation assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Proteome Coverage
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Pathway Enrichment Consistency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Reproducibility Information Completeness
The extent to which all scientifically necessary components of reproducibility information are present, documented, and evaluable.
bemo
BEMO:2000430
BEMO:2000430
437
3aaeec4ac26a2a3b26b615c686a2f19768e2153bd9e1e676bdcd1f2ba2812fc8
10
Assesses reproducibility information completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in reproducibility information completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reproducibility information completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Sample Size Justification
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Hardy–Weinberg Equilibrium Compatibility
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of hardy–weinberg equilibrium compatibility.
bemo
BEMO:2000248
BEMO:2000248
255
158d6abc0d325612b86f33eeb0aa8e0dbdfd4460dc9c1934f90601f28ed25284
8
Assesses hardy–weinberg equilibrium compatibility using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in hardy–weinberg equilibrium compatibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating hardy–weinberg equilibrium compatibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Verification Bias Risk
The probability or degree that verification bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000138
BEMO:2000138
145
2f165758b70b1e482bfb8517e1a69a7d4e05ba84cd41dc36c2f346f991371034
10
Assesses verification bias risk using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in verification bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating verification bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Metabolic Feature Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Dynamic Range Coverage
The proportion and representativeness of the relevant dynamic range captured by the evidence or measurement process.
bemo
BEMO:2000361
BEMO:2000361
368
57e4c77ea2a48565f0758aadf8529c7db52bab9beb67f6e9c8366f0f3232e753
9
Assesses dynamic range coverage using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in dynamic range coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating dynamic range coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
RNA Evidence Strength
The magnitude and credibility of independent evidence supporting rna evidence.
bemo
BEMO:2000231
BEMO:2000231
238
8e776a34bb92ed3d093c7eb03481f1c61ce724a0995235355586c28e53418e35
8
Assesses rna evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in rna evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating rna evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Context Specific Metric Computation
Controlled BEMO ComputationMode value: ContextSpecificMetricComputation.
BEMO:4000007
ContextSpecificMetricComputation
Computation profile for Multiplicity-Adjusted Credibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Trueness
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Threshold Validity
The degree to which threshold supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000136
BEMO:2000136
143
c27bfe6217acf4b748aeb85aa4e9c4176c8256e12b7c61ff84a201b6fd374c82
10
Assesses threshold validity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in threshold validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating threshold validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Species Appropriateness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Species Extrapolation Validity
The degree to which species extrapolation supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000354
BEMO:2000354
361
f3eb077ab3d21773df18d06eb29c8a11272ed7cbfe8f22da0ffedd8a171bc393
10
Assesses species extrapolation validity using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in species extrapolation validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating species extrapolation validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Material Availability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of material availability.
bemo
BEMO:2000398
BEMO:2000398
405
ad24de4b8dd0c9ae8c1a3d7d5fcdd166b9bd923967cbf8032882e26fce35feb2
8
Assesses material availability using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in material availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating material availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Evidence Consistency
The degree of agreement in evidence across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000146
BEMO:2000146
153
f74a279906eee7e4453803727494ade321b94fc1d6196fb1882daae7f1abcf75
8
Assesses evidence consistency using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Cutoff Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Negative-Result Reporting
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Real-World Evidence Alignment
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of real-world evidence alignment.
bemo
BEMO:2000204
BEMO:2000204
211
2bcc973aa3f12efded260e2aef1541636262ee110856a34255ce0823d91dfb20
8
Assesses real-world evidence alignment using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in real-world evidence alignment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating real-world evidence alignment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Biospecimen and Preanalytical Quality metric
Category of biomedical evidence metrics concerned with biospecimen and preanalytical quality.
bemo
BEMO:1100003
Candidate
Structural Identifiability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of structural identifiability.
bemo
BEMO:2000328
BEMO:2000328
335
36fbdf592582d2c048b53d43d95edfca8afe5d6c7c93d352f32f2a7d7644cfbe
8
Assesses structural identifiability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in structural identifiability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating structural identifiability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Differential Expression Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of differential expression robustness.
bemo
BEMO:2000244
BEMO:2000244
251
e70802b2991914211c0dcd91213c02319dd673fc93963b9e2064b8bca944c528
8
Assesses differential expression robustness using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in differential expression robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating differential expression robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Exchangeability Plausibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Epistasis Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Detection Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Dose–Response Support
0.1.0
Apply a validated analyte- and method-specific measurement procedure with calibration and quality control.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","unitRef":"REQUIRED"}
specimen_or_material; measurement_procedure; calibration_reference; quality_control_results; unit
replicate_measurements; environmental_conditions; instrument_version
xsd:decimal
Method- and analyte-specific physical units
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Allocation Concealment
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Experimental Reproducibility
The degree to which experimental reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000393
BEMO:2000393
400
6c9d1f649c85d73c0009300283cb79e817b0ca9b1d53444f1a4aa4b5a2e6d8ae
10
Assesses experimental reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in experimental reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating experimental reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Residual Confounding Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
No-Interference Plausibility
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of no-interference plausibility.
bemo
BEMO:2000101
BEMO:2000101
108
8aa9a1ba76fa5fd8217fff1b4352656d262602a1040dc8618b138b2e37a93d9a
9
Assesses no-interference plausibility using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in no-interference plausibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating no-interference plausibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 217
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Sampling Frame Adequacy; Generalizability; Setting Applicability", "Common Misinterpretations": "Treating transportability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Transportability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of transportability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses transportability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in transportability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
217
ef0b46b07942a1780f9e2e890b09b7edc3579f016762848aa46c4d956da3075d
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 282
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Cross-Reactivity; Hook Effect Risk; Calibration Traceability", "Common Misinterpretations": "Treating carryover as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Carryover", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of carryover.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses carryover using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in carryover can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
282
0e987b7e93f5f6b586140ce3b08eb3d87acd863bc4cb825ee0d1e4b7b1bb3994
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 261
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Strand Bias; Call-Rate Completeness; Hardy–Weinberg Equilibrium Compatibility", "Common Misinterpretations": "Treating reference bias as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Reference Bias", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of reference bias.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses reference bias using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reference bias can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
261
7f170b7ac0b632bb15937074f6f54297bd1374274ba7dd6a71c308a1270c144b
Quantification Accuracy
The closeness of quantification to the accepted reference or true value.
bemo
BEMO:2000381
BEMO:2000381
388
5205abd1da59c18e5b7272ce8586c2ec2bee590ad36aa3289ce71f4aa4c17c0b
8
Assesses quantification accuracy using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in quantification accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating quantification accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Subgroup Consistency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Diagnostic Specificity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Time-Varying Confounding Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Trueness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of trueness.
bemo
BEMO:2000302
BEMO:2000302
309
9094cf98fb975c34384185560acacc946dab62ebd8f1815100dcaac099d9df1b
8
Assesses trueness using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in trueness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating trueness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Model Fit
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of model fit.
bemo
BEMO:2000449
BEMO:2000449
456
bf6f70bdb1883a9c1ba70dc7d37eb1885dd46ed70094a2b12bd2cc41e0f53bad
8
Assesses model fit using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in model fit can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating model fit as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Overadjustment Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Small-Study Effects
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Animal Model Construct Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 438
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Harms Reporting Completeness; Missing-Data Reporting Completeness; Funding-Source Transparency", "Common Misinterpretations": "Treating statistical methods reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Statistical Methods Reporting Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of statistical methods reporting are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses statistical methods reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in statistical methods reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
438
d8af14618f3e80ecff3b617a35c54caa9b3c73184f234fcea88a9ee83c7090ff
uncertainty specification
An information content entity specifying how uncertainty is estimated, propagated, calibrated, and reported.
bemo
BEMO:0000106
Candidate
Computation profile for Reverse-Causation Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Imputation Validity
The degree to which imputation supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000443
BEMO:2000443
450
451bb1f13aa8cd7feb728a231f72c948355ace1a73de7f1903381352517ceaf5
10
Assesses imputation validity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in imputation validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating imputation validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Computation profile for Dilution Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Batch-Effect Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Individual-Level Surrogacy
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of individual-level surrogacy.
bemo
BEMO:2000045
BEMO:2000045
52
4d46697f970727d741077ee1e1f9eeeaed4beee8d6ef96a339ca927181b91ef5
8
Assesses individual-level surrogacy using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in individual-level surrogacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating individual-level surrogacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Loss-of-Function Validation
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Phenocopy Risk
The probability or degree that phenocopy introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000228
BEMO:2000228
235
eba377fe2c190b6390cd804f64874ad58540debe287184cfa259e70d97d1fcc8
8
Assesses phenocopy risk using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in phenocopy risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating phenocopy risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Spatial Transcriptomic Registration Accuracy
The closeness of spatial transcriptomic registration to the accepted reference or true value.
bemo
BEMO:2000263
BEMO:2000263
270
e61d8e9d42dcb8b644bc33652bdf4f396f77ace1c2d848fb684d0c722e5a1598
8
Assesses spatial transcriptomic registration accuracy using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in spatial transcriptomic registration accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating spatial transcriptomic registration accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Animal Model Predictive Validity
The degree to which animal model predictive supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000169
BEMO:2000169
176
ae7f68af5b958c01b88066cf76664c77d333abaff4414b63760b817bf19359d0
10
Assesses animal model predictive validity using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in animal model predictive validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating animal model predictive validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Fragility Index
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of fragility index.
bemo
BEMO:2000442
BEMO:2000442
449
045c39ed086cc31d598f1fe91547a90170ea8cc0b684f22d5338144e750a0843
8
Assesses fragility index using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in fragility index can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating fragility index as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Blinding Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Selection-on-Survival Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Splicing Evidence Strength
The magnitude and credibility of independent evidence supporting splicing evidence.
bemo
BEMO:2000233
BEMO:2000233
240
50c6384810a1443b5a98f42466fb8ab2489b57dad29ed2dab35ef7acec1960d2
8
Assesses splicing evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in splicing evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating splicing evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Storage Temperature Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Outcome Applicability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of outcome applicability.
bemo
BEMO:2000201
BEMO:2000201
208
5c0db0a0257f7b171655d97624fdf4764d1fabad77716300abc92c2debe71f7d
8
Assesses outcome applicability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in outcome applicability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating outcome applicability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Falsification Test Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computational Reproducibility
The degree to which computational reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000388
BEMO:2000388
395
c15f070192e578e50deb91384b3e8099e5c4b1a0c28cc84d1ac6a3163d36cd17
10
Assesses computational reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in computational reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating computational reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 186
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Intervention Fidelity in Animal Studies; Attrition Accounting in Animal Studies; Exclusion-Criteria Prespecification", "Common Misinterpretations": "Treating humane endpoint appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Humane Endpoint Appropriateness", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The extent to which humane endpoint is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses humane endpoint appropriateness using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in humane endpoint appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
186
efafe26c06dac15c0d485bfb264933c40fe9cf7ca97c5195ea92c9cfeb3f7a98
Contamination Burden
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of contamination burden.
bemo
BEMO:2000242
BEMO:2000242
249
fc0cd98bc92235d469bf2bc025f0d327d8170dfb4df97b973710a8d887eff3f4
8
Assesses contamination burden using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in contamination burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating contamination burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Evidence Coverage
The proportion and representativeness of the relevant evidence captured by the evidence or measurement process.
bemo
BEMO:2000147
BEMO:2000147
154
a64cee9a2114ea7d5580ea8d59eaef0691fce1d63dd3e4fd9f3d63e0dac86a3d
9
Assesses evidence coverage using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 397
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Code Availability; Random-Seed Stability; Researcher-Degrees-of-Freedom Sensitivity", "Common Misinterpretations": "Treating data availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Data Availability", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of data availability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses data availability using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in data availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
397
89d3d59a31fe591c17a3c61f83a55d18326e4b311b8033047e95bfb2a93715fa
Computation profile for Prespecified Analysis Adherence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 49
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Outcome Relevance; Endpoint Responsiveness; Minimal Clinically Important Difference Validity", "Common Misinterpretations": "Treating endpoint reliability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Endpoint Reliability", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of endpoint reliability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses endpoint reliability using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in endpoint reliability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
49
e147bf4116de09d3c1bd6b5fb8081223589d31f54f697b6e869790830619757d
Study Heterogeneity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of study heterogeneity.
bemo
BEMO:2000165
BEMO:2000165
172
93fce0e52f596e9f5e6b6636a82da697803846e686e6144faa894115a270196a
8
Assesses study heterogeneity using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in study heterogeneity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Forest plots; heterogeneity statistics; tau-squared; prediction intervals; funnel plots; regression or selection models; sensitivity analyses.
Treating study heterogeneity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Statistical Power
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of statistical power.
bemo
BEMO:2000459
BEMO:2000459
466
b5866270dc0b1ae7d14dc9e42f0ff6a1d28896f404d73ddeb44c5f7fdaf6b51a
8
Assesses statistical power using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in statistical power can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating statistical power as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Performance Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Housing and Husbandry Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of housing and husbandry control.
bemo
BEMO:2000178
BEMO:2000178
185
d4c26e56573de7f55de0caafd1e2880580fe66b8b4d599787ebf0b2bb3808344
8
Assesses housing and husbandry control using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in housing and husbandry control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating housing and husbandry control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Hardy–Weinberg Equilibrium Compatibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Researcher-Degrees-of-Freedom Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Test-Timing Appropriateness
The extent to which test-timing is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000135
BEMO:2000135
142
ffb2707e3bed89aca5f5c790dc5a9b0312e89c5a1dba27ebac908e4597a355bd
8
Assesses test-timing appropriateness using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in test-timing appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating test-timing appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Analytical Precision
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Multiplicity Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Intervention Description Completeness
The extent to which all scientifically necessary components of intervention description are present, documented, and evaluable.
bemo
BEMO:2000415
BEMO:2000415
422
52ff9263e3651129913da102dae69c16126528e71dde3bf30595ca84301723b5
9
Assesses intervention description completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in intervention description completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating intervention description completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Ion Suppression Assessment
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of ion suppression assessment.
bemo
BEMO:2000366
BEMO:2000366
373
c108ebf1ca903ed7e96063f24ecaf39723dd5e925f97d641601f89fd859897b9
8
Assesses ion suppression assessment using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in ion suppression assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating ion suppression assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Eligibility Criteria Completeness
The extent to which all scientifically necessary components of eligibility criteria are present, documented, and evaluable.
bemo
BEMO:2000412
BEMO:2000412
419
02e23f40dd856c6f19f825d6191af4cb70f9ed2e604c47f5cbd7bf406e4d2381
9
Assesses eligibility criteria completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in eligibility criteria completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating eligibility criteria completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Hook Effect Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Prognostic Discrimination
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prognostic discrimination.
bemo
BEMO:2000131
BEMO:2000131
138
b58deaf6b24f17f45025c2e10bd742dcb00a937cc0410366fbe092bfbfd13882
8
Assesses prognostic discrimination using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in prognostic discrimination can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating prognostic discrimination as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Isotope Pattern Fidelity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Convergent Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Locus Heterogeneity Assessment
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of locus heterogeneity assessment.
bemo
BEMO:2000224
BEMO:2000224
231
bd3af304ec15826df5f67db7f079c0ba66a1f5f5a3e852d600550186267f50cb
8
Assesses locus heterogeneity assessment using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in locus heterogeneity assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Forest plots; heterogeneity statistics; tau-squared; prediction intervals; funnel plots; regression or selection models; sensitivity analyses.
Treating locus heterogeneity assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Estimate Precision
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Efficacy Reproducibility
The degree to which efficacy reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000336
BEMO:2000336
343
e3d46320b779b7100a656dae6dcb02f7b9eae6762c1680f48b9313843cb64c84
10
Assesses efficacy reproducibility using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in efficacy reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating efficacy reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Conceptual Replication Success
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Metabolite Identification Confidence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Emergent-Property Reproducibility
The degree to which emergent-property reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000307
BEMO:2000307
314
e554ab3ba02f294f767076d372c3967762b850eccc356bf4d7c243480641cd98
10
Assesses emergent-property reproducibility using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in emergent-property reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating emergent-property reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Specialized / infrequent
Developing
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 383
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Fragmentation Spectrum Quality; Proteoform Identification Confidence; Metabolite Identification Confidence", "Common Misinterpretations": "Treating post-translational modification localization confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Post-Translational Modification Localization Confidence", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The justified degree of certainty assigned to post-translational modification localization given the quantity, quality, consistency, and limitations of supporting evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses post-translational modification localization confidence using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in post-translational modification localization confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
383
6761ab89a934b85357a2d58a5e53503065037a8cd658104babdb5536291a3653
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 331
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Flux-Balance Consistency; Emergent-Property Reproducibility", "Common Misinterpretations": "Treating perturbation prediction accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Perturbation Prediction Accuracy", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The closeness of perturbation prediction to the accepted reference or true value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses perturbation prediction accuracy using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in perturbation prediction accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
331
5cc466b3de703eb697570667c57a9230690c56d69b5cf6ec528e42740ec621e9
Computation profile for Spectral Library Match Quality
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 475
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Co-intervention Bias Risk; Period Effect Risk; Cluster Recruitment Bias Risk", "Common Misinterpretations": "Treating carryover effect risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Carryover Effect Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that carryover effect introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses carryover effect risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in carryover effect risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
475
6dbe1d75153c55efcce7da9af8dda3fe869c2f426eaaad651ee6f91b1670d205
Computation profile for Interaction Assessment Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Commutability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Inferential Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Co-intervention Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Calibration-in-the-Large
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Exposure Classification Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Equivalence Margin Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Surrogate Endpoint Validity
The degree to which surrogate endpoint supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000055
BEMO:2000055
62
880b1c4a9fa2a25e2876679a2a7e759365591c8d0d0ae431e9a6e89a9d327bd9
10
Assesses surrogate endpoint validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in surrogate endpoint validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating surrogate endpoint validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Collider Bias Risk
The probability or degree that collider bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000087
BEMO:2000087
94
0d0ee23715a02b538d00c3772dc0b2a1a64aba9fb6a21ecd62845d74c88b63ca
10
Assesses collider bias risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in collider bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating collider bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 64
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Individual-Level Surrogacy; Endpoint Validity; Outcome Relevance", "Common Misinterpretations": "Treating trial-level surrogacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Trial-Level Surrogacy", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of trial-level surrogacy.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses trial-level surrogacy using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in trial-level surrogacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
64
977d5f2b7e02b5e90f442ea4e8788e2fe539316cddda3ae61da29b0fa71c5325
Intermediate Precision
The closeness of repeated estimates or measurements and the narrowness of uncertainty around intermediate.
bemo
BEMO:2000284
BEMO:2000284
291
543dc6e95f52dce066f892ad8d141154500ea9d0cf35d11564c84cc551f70c3b
9
Assesses intermediate precision using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in intermediate precision can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating intermediate precision as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Cluster Recruitment Bias Risk
The probability or degree that cluster recruitment bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000469
BEMO:2000469
476
a9882065e899d9f8a8f701c6291666b702a9ad5b769971c1d3cb596c9e39b960
10
Assesses cluster recruitment bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in cluster recruitment bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cluster recruitment bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Sequencing Depth Adequacy
The extent to which sequencing depth is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000257
BEMO:2000257
264
719d5f0d3157f75114503dd8e8b0e75aa481929ab009d6c25b68024c7b5f71a0
8
Assesses sequencing depth adequacy using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in sequencing depth adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.
Treating sequencing depth adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
evidence snapshot identifier
Identifier of the immutable evidence snapshot.
BEMO:3100029
Computation profile for Technical Replicate Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Time-to-Fixation Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 492
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Allocation Concealment; Baseline Comparability", "Common Misinterpretations": "Treating randomization integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Randomization Integrity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of randomization integrity.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses randomization integrity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in randomization integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
492
8438f239c984919f366a44747e2e32efd683c0e45e3f7af6c269801f1d6694c8
Computation profile for Biological Gradient
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Reference Bias
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Geographic Consistency
The degree of agreement in geographic across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000199
BEMO:2000199
206
22cb03e816656f6f9ad7a2acb4c32a01ffbc6b40d04a48011cef0a52134d504e
8
Assesses geographic consistency using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in geographic consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating geographic consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 146
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Study Heterogeneity; Prediction Interval Adequacy; Cumulative Evidence Stability", "Common Misinterpretations": "Treating between-study variance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Between-Study Variance", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of between-study variance.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Forest plots; heterogeneity statistics; tau-squared; prediction intervals; funnel plots; regression or selection models; sensitivity analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses between-study variance using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in between-study variance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
146
167d33dabeef5837833710d7e97d42ff17e46ede74b98e0e88db46e921ff811e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 354
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Exposure–Response Relationship; Pharmacodynamic Adequacy; Bioavailability", "Common Misinterpretations": "Treating pharmacokinetic adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Pharmacokinetic Adequacy", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The extent to which pharmacokinetic is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses pharmacokinetic adequacy using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in pharmacokinetic adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
354
2b02efdd0f6a774096cee910d90b575b8f28f136a8ee594876642fb213f66330
Library Complexity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of library complexity.
bemo
BEMO:2000249
BEMO:2000249
256
fb650cc4b9fc46d550895b59d1656dada1a7ec625426686e0a708ba75b6fc5f0
8
Assesses library complexity using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in library complexity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating library complexity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Ontology Annotation Completeness
The extent to which all scientifically necessary components of ontology annotation are present, documented, and evaluable.
bemo
BEMO:2000318
BEMO:2000318
325
b8714924720c209e415f9bf2c9aaf707f6bf0edbbbf6ca8da11d2546f9da563d
9
Assesses ontology annotation completeness using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in ontology annotation completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating ontology annotation completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Harms Reporting Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Statistical Power
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 42
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Criterion Validity; Content Validity; Predictive Biomarker Validity", "Common Misinterpretations": "Treating construct validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Construct Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which construct supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses construct validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in construct validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
42
5bd558ed006f260bde44545e9ab38951fc6c2e450b4d3cb07f929d1fcb479f4e
Computation profile for Tissue Specificity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Single-Cell Doublet Burden
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of single-cell doublet burden.
bemo
BEMO:2000260
BEMO:2000260
267
47e6d4bd3f64b35d8abb9925b05e112fbb1356d383dc641a47f78f8e9bbeb676
8
Assesses single-cell doublet burden using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in single-cell doublet burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating single-cell doublet burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Immortal-Time Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Competing-Risk Model Validity
The probability or degree that competing-risk model validity introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000114
BEMO:2000114
121
6f322698057cc9561b413498181b8f2503720420a935d144d2173ae5298a3b2f
10
Assesses competing-risk model validity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in competing-risk model validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating competing-risk model validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Index-Test Blinding
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Positive Likelihood Ratio
0.1.0
value = numerator / denominator; the null value is typically 1 where scientifically applicable
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator != 0"],"null_value":1}
numerator; denominator; operational_definition; assessment_context
confidence_level; stratum
xsd:decimal
Ratio scale; null typically 1
0.0
1
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Not generally required unless converted to a probability or score.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Cluster Recruitment Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Protocol Fidelity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Cross-Layer Directional Concordance
The degree of agreement in cross-layer directional across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000303
BEMO:2000303
310
89f675805a18a02e8214d03853d6a7cdcea2d29079ee839babae7605cea7a581
8
Assesses cross-layer directional concordance using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in cross-layer directional concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cross-layer directional concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Protein Inference Reliability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of protein inference reliability.
bemo
BEMO:2000378
BEMO:2000378
385
601ad0978c475da6065d2502b00169f5521ae97cb0543ad15ad88f49903cb27e
8
Assesses protein inference reliability using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in protein inference reliability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating protein inference reliability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Inferential Reproducibility
The degree to which inferential reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000395
BEMO:2000395
402
b49eb32e7bf118727039f0240b17aee8970c665e4a359a2bcecc87341188ecf1
10
Assesses inferential reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in inferential reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating inferential reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Call-Rate Completeness
The extent to which all scientifically necessary components of call-rate are present, documented, and evaluable.
bemo
BEMO:2000240
BEMO:2000240
247
c69a25da4615faa5675944ca32d4c5ec30a0c44f5216034581329797d9a32abc
9
Assesses call-rate completeness using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in call-rate completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.
Treating call-rate completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 126
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Differential Verification Bias Risk; Patient Flow Integrity; Test-Timing Appropriateness", "Common Misinterpretations": "Treating incorporation bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Incorporation Bias Risk", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "The probability or degree that incorporation bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses incorporation bias risk using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in incorporation bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
126
00ceeb2a9486ba76c68ad0f07ad3fe11fadf573e3b3b3fb37898281142b34a91
Selection-on-Survival Bias Risk
The probability or degree that selection-on-survival bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000106
BEMO:2000106
113
ba5e2bde29a386ccb9cae3afe1f1c9476a350f65371bc1aebd58a3579e6e702c
10
Assesses selection-on-survival bias risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in selection-on-survival bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating selection-on-survival bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Nonlinearity Assessment
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of nonlinearity assessment.
bemo
BEMO:2000453
BEMO:2000453
460
aec401720de6539460ff9d49f62284a772b28790802eda1b3dc7c5fc2d67af3f
8
Assesses nonlinearity assessment using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in nonlinearity assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating nonlinearity assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Extraction Recovery
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of extraction recovery.
bemo
BEMO:2000362
BEMO:2000362
369
5024c9034e231e8bc296cf1d9092cc063fe94890f4856d8bf70ac077514b55f9
8
Assesses extraction recovery using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in extraction recovery can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating extraction recovery as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Data Availability
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Protocol Deviation Risk
The probability or degree that protocol deviation introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000483
BEMO:2000483
490
7632dd2b67a9f407b6c186b3a54ef8714ff7082ed9c885957d4f06c0008e5f26
8
Assesses protocol deviation risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in protocol deviation risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating protocol deviation risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Relation Evidence Strength
The magnitude and credibility of independent evidence supporting relation evidence.
bemo
BEMO:2000326
BEMO:2000326
333
9009911ac3f6266273d49aa664b528314dd25468ab2969a5c3dfb8593502a99c
8
Assesses relation evidence strength using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in relation evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating relation evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Prior Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Analytical Specificity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Negative-Control Validation
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of negative-control validation.
bemo
BEMO:2000100
BEMO:2000100
107
63bedc3fc09ef48160a8d522d76618e50e47197d3d502ae109f1015d0fbcb67e
8
Assesses negative-control validation using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in negative-control validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating negative-control validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Genetic Background Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Carryover Effect Risk
The probability or degree that carryover effect introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000468
BEMO:2000468
475
6dbe1d75153c55efcce7da9af8dda3fe869c2f426eaaad651ee6f91b1670d205
8
Assesses carryover effect risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in carryover effect risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating carryover effect risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Setting Applicability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Material Availability
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Protein Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of protein integrity.
bemo
BEMO:2000077
BEMO:2000077
84
f776649240f851879898d488bb90516491936e0936f6f03ac5946be6c8b62a20
8
Assesses protein integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in protein integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating protein integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Normalization Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Relation Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Code-Sharing Transparency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Carcinogenicity Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Influential Observation Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of influential observation sensitivity.
bemo
BEMO:2000444
BEMO:2000444
451
0202d30b443e9ba2dd8ad4597c8e7b9fcc240af0f60b19db4a115b1d2866e858
9
Assesses influential observation sensitivity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in influential observation sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating influential observation sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Computation profile for Allelic Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Evidence Directness
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Parameter Identifiability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of parameter identifiability.
bemo
BEMO:2000320
BEMO:2000320
327
e1eaab072c302dca725ea1fb92e565c0aead2c5a6ddd49490dbfa43ad163dd03
8
Assesses parameter identifiability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in parameter identifiability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating parameter identifiability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Reanalysis Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 195
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Orthogonal Validation", "Common Misinterpretations": "Treating technical artifact exclusion as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Technical Artifact Exclusion", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of technical artifact exclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses technical artifact exclusion using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in technical artifact exclusion can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
195
8d21017136fa388d87b531d9b552fc3f6787a0b2615866e0c0a1ca9ae2fb0bbf
Generalizability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of generalizability.
bemo
BEMO:2000198
BEMO:2000198
205
27f8a61447192cbdeaaa7e84078183a3604b2e6ac6c0df13be6cfe9e4a09224a
8
Assesses generalizability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in generalizability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating generalizability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Allelic Balance
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of allelic balance.
bemo
BEMO:2000237
BEMO:2000237
244
0857b2292263da43e4b892ac8b30335b2a7fd8730d0baae8269f0919f1da549b
8
Assesses allelic balance using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in allelic balance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating allelic balance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 320
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Ontology Evidence-Code Strength; Parameter Identifiability; Parameter Sensitivity", "Common Misinterpretations": "Treating model–experiment concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Model–Experiment Concordance", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The degree of agreement in model–experiment across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses model–experiment concordance using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in model–experiment concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
320
433f97851837c60d123f8a332258d7d5b3872aa7c5564c19691b4c4c00871a8b
Computation profile for Variance Estimation Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 216
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Geographic Consistency; Demographic Generalizability; Disease-Severity Generalizability", "Common Misinterpretations": "Treating temporal generalizability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Temporal Generalizability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of temporal generalizability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses temporal generalizability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in temporal generalizability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
216
dc35bc0fc74498ddd3acc1a47faf52ec96fba459b80f9800025d1a3d1e674317
Exchangeability Plausibility
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of exchangeability plausibility.
bemo
BEMO:2000095
BEMO:2000095
102
7f4004986e0cd5fd0b8462346aec167d3be6e60fe67bb8a5838de7a005fd58d1
9
Assesses exchangeability plausibility using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in exchangeability plausibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.
Treating exchangeability plausibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Safety Margin Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for E-value Strength
0.1.0
value = numerator / denominator; the null value is typically 1 where scientifically applicable
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator != 0"],"null_value":1}
numerator; denominator; operational_definition; assessment_context
confidence_level; stratum
xsd:decimal
Ratio scale; null typically 1
0.0
1
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Not generally required unless converted to a probability or score.
Source maturity: Developing; BEMO computation profile requires independent validation.
Alternative Splicing Validation
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of alternative splicing validation.
bemo
BEMO:2000238
BEMO:2000238
245
64b27913521868b5a9491d569eb82729fee0c3cb2133bfc9132e110785fa8847
8
Assesses alternative splicing validation using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in alternative splicing validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating alternative splicing validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Prognostic Discrimination
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Consistency Assumption Plausibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Operator Variability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Randomization Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of randomization integrity.
bemo
BEMO:2000485
BEMO:2000485
492
8438f239c984919f366a44747e2e32efd683c0e45e3f7af6c269801f1d6694c8
10
Assesses randomization integrity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in randomization integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating randomization integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Computation profile for Blinding in Experimental Assessment
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Missing-Value Burden
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 224
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Segregation Evidence Strength; Allelic Evidence Strength; Case-Level Evidence Strength", "Common Misinterpretations": "Treating de novo evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "De Novo Evidence Strength", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The magnitude and credibility of independent evidence supporting de novo evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses de novo evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in de novo evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
224
e1a17a41d450c91691b405a890d0fa9af7ecd8fed12385f8124d85b912c16860
Computation profile for Limit of Detection
0.1.0
Apply a validated analyte- and method-specific measurement procedure with calibration and quality control.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","unitRef":"REQUIRED"}
specimen_or_material; measurement_procedure; calibration_reference; quality_control_results; unit
replicate_measurements; environmental_conditions; instrument_version
xsd:decimal
Method- and analyte-specific physical units
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required.
Source maturity: Mature; BEMO computation profile requires independent validation.
Intervention Applicability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of intervention applicability.
bemo
BEMO:2000200
BEMO:2000200
207
d1cb3929ecc9b0976fba4823cb5a09ca13a634dd6efaa3ee43969eafe3250a95
8
Assesses intervention applicability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in intervention applicability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating intervention applicability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Individual-Level Surrogacy
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Missing-Value Burden
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of missing-value burden.
bemo
BEMO:2000372
BEMO:2000372
379
770bd2fa8c1eedf14831e8e8caca6d53a1ad4a7b7272433891be0eac41a4e30c
8
Assesses missing-value burden using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in missing-value burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating missing-value burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Outcome Relevance
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of outcome relevance.
bemo
BEMO:2000049
BEMO:2000049
56
1adca583ee04c245cdae3428bd050f3ed8796d29bf4b214b5880769da8977172
8
Assesses outcome relevance using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in outcome relevance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating outcome relevance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Ecological Validity
The degree to which ecological supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000196
BEMO:2000196
203
04e73707dc619457f6718e0be03029ef6e99208f68cbcad30d9f5bb20b2cd94b
10
Assesses ecological validity using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in ecological validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating ecological validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Off-Target Liability Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Intervention Applicability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Prognostic Calibration
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prognostic calibration.
bemo
BEMO:2000130
BEMO:2000130
137
bddeb4ff61529649f6f5d81ae9a9ab5588a6f6aed10ca70d0d400119d6c2e940
8
Assesses prognostic calibration using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in prognostic calibration can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating prognostic calibration as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Outcome Applicability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Falsification Test Support
The magnitude and credibility of independent evidence supporting falsification test.
bemo
BEMO:2000096
BEMO:2000096
103
280760fd65dbda26ef711b182bed36c1b7e3679af41be52e50dfca26f92b3400
8
Assesses falsification test support using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in falsification test support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating falsification test support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Comparator Validity
The degree to which comparator supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000471
BEMO:2000471
478
de06a641ed9f64a7c11259e9a1d1346140c2469d57bbd106c6257d947909a5a3
10
Assesses comparator validity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in comparator validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating comparator validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Computation profile for Direct Replication Success
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Mechanistic Support
The magnitude and credibility of independent evidence supporting mechanistic.
bemo
BEMO:2000014
BEMO:2000014
21
d05eb179325de1077bb0eedc4d22a5f8c9cc2049d4ca5a5830b380984606cc90
8
Assesses mechanistic support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in mechanistic support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating mechanistic support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Patient Flow Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Computational Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Researcher-Degrees-of-Freedom Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of researcher-degrees-of-freedom sensitivity.
bemo
BEMO:2000404
BEMO:2000404
411
c196ee0e5b3c49cffc78d559a15b400e1dec74f4854200e7fbd341eef60a59c4
9
Assesses researcher-degrees-of-freedom sensitivity using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in researcher-degrees-of-freedom sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating researcher-degrees-of-freedom sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 245
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Transcript Quantification Reliability; Single-Cell Doublet Burden; Single-Cell Ambient RNA Burden", "Common Misinterpretations": "Treating alternative splicing validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Alternative Splicing Validation", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of alternative splicing validation.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses alternative splicing validation using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in alternative splicing validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
245
64b27913521868b5a9491d569eb82729fee0c3cb2133bfc9132e110785fa8847
Computation profile for Case-Level Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Unmeasured Confounding Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Retention-Time Stability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of retention-time stability.
bemo
BEMO:2000383
BEMO:2000383
390
7ed489b83705dc957b891fd6483e7b54ca4c70547f903f7dfa9e53d0ca59fa58
9
Assesses retention-time stability using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in retention-time stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating retention-time stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Protein Identification Confidence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Genotype–Phenotype Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Residual Confounding Risk
The probability or degree that residual confounding introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000104
BEMO:2000104
111
fef3c7d8a6ff0d0d4088a08abb2cfe7420127e6170b4998b7b5ce06d44124e81
10
Assesses residual confounding risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in residual confounding risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.
Treating residual confounding risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Library Complexity
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Postanalytical Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pathway Topology Support
The magnitude and credibility of independent evidence supporting pathway topology.
bemo
BEMO:2000323
BEMO:2000323
330
f408d8b817c9e3179698dad395d8e94ea187b92ae192e0ab52e026106d82f40a
8
Assesses pathway topology support using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in pathway topology support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating pathway topology support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Phenotypic Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
RNA Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of rna integrity.
bemo
BEMO:2000078
BEMO:2000078
85
98ac509f170b3163d862e8e08cc16dad8b915df5508f4e7394a8627d881f7350
8
Assesses rna integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in rna integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating rna integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pathway Enrichment Consistency
The degree of agreement in pathway enrichment across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000322
BEMO:2000322
329
d7e323b8eddcaff29fe662695a2fa9bae2a0f04120cfdf422fa3e2d1cd3b117e
8
Assesses pathway enrichment consistency using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in pathway enrichment consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating pathway enrichment consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Reproductive Toxicity Evidence Strength
The magnitude and credibility of independent evidence supporting reproductive toxicity evidence.
bemo
BEMO:2000351
BEMO:2000351
358
1725f0845040d88689ff2869cde10ea44454b5cfaa1c303d02071f7ff5ccb751
10
Assesses reproductive toxicity evidence strength using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in reproductive toxicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reproductive toxicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Co-segregation Likelihood
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of co-segregation likelihood.
bemo
BEMO:2000214
BEMO:2000214
221
965b0273f9a305e64b50437dcad85ab39818fad51d561d0ac4da1daaf7d204f2
8
Assesses co-segregation likelihood using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in co-segregation likelihood can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating co-segregation likelihood as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Time-Dependent Discrimination
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Early Stopping Bias Risk
The probability or degree that early stopping bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000476
BEMO:2000476
483
8ac86a4a6b6a886dad127939918a466b8c70bd2b6b13b03a987aab048ac4c7c3
10
Assesses early stopping bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in early stopping bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating early stopping bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Computation profile for Immunotoxicity Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 19
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Mechanistic Support; Mechanistic Completeness; Mechanistic Coherence", "Common Misinterpretations": "Treating mechanistic coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Mechanistic Coverage", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The proportion and representativeness of the relevant mechanistic captured by the evidence or measurement process.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses mechanistic coverage using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in mechanistic coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
19
8582aac321a237f1dd1db68f0893021d3bc987b71b373d1879cc2dfb098adbdb
Computation profile for Quality-Control Reporting Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Tumor Purity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of tumor purity.
bemo
BEMO:2000082
BEMO:2000082
89
e7836bb1f6728d199c8634dbdbf560dae6243a3932920cc82164935a37c9180e
8
Assesses tumor purity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in tumor purity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating tumor purity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Biomarker Reliability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
aggregation rule
Aggregation rule for multi-input or composite computations.
BEMO:3100013
Potency Reproducibility
The degree to which potency reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000349
BEMO:2000349
356
2199b92ac6c23ce1866fa6fdae782c65a478ec51834e2f0193f79d608ea4c3e7
10
Assesses potency reproducibility using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in potency reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating potency reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Spatial Biological Concordance
The degree of agreement in spatial biological across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000023
BEMO:2000023
30
6c08ba26b1257f37b62212d7633a9e83166a0f74706ae3ffbb9de90cd98613f7
8
Assesses spatial biological concordance using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in spatial biological concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating spatial biological concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Correct Temporal Ordering
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of correct temporal ordering.
bemo
BEMO:2000090
BEMO:2000090
97
b11d505f3985b03d1ab78fe50a79558d187aeacc3fabbc39b5dc3da61a0cfbe2
8
Assesses correct temporal ordering using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in correct temporal ordering can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating correct temporal ordering as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Randomization in Experimental Allocation
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Sample Stability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
source evidence level
Preserves the source evidence-level scope.
BEMO:3200008
Biospecimen Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biospecimen integrity.
bemo
BEMO:2000059
BEMO:2000059
66
b89a21b15f8ca110af0a2e2a28ae4ce7939819f7d62769c459c1cd5dc9b712cc
8
Assesses biospecimen integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in biospecimen integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biospecimen integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Diagnostic Biomarker Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Linearity
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 386
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Post-Translational Modification Localization Confidence; Metabolite Identification Confidence; Metabolite Annotation Level", "Common Misinterpretations": "Treating proteoform identification confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Proteoform Identification Confidence", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The justified degree of certainty assigned to proteoform identification given the quantity, quality, consistency, and limitations of supporting evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses proteoform identification confidence using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in proteoform identification confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
386
8df86e86625f432ba76b927808300a9cfa4a9b6844cae40ca0c6240f4f899eda
Computation profile for Reference Standard Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Null-Variant Quality
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 447
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Prior Sensitivity; Noninferiority Margin Validity", "Common Misinterpretations": "Treating equivalence margin validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Equivalence Margin Validity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The degree to which equivalence margin supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses equivalence margin validity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in equivalence margin validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
447
226b21fe561d36c12bfde0bcc47bb1ad6cb02760533a4b768747d6a0ab92184b
Computation profile for Positivity Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Rescue Experiment Support
The magnitude and credibility of independent evidence supporting rescue experiment.
bemo
BEMO:2000022
BEMO:2000022
29
9f3de30fb4252544b664477f26ad879c47cffd18e4a902652a07ca5b9fa59ec2
8
Assesses rescue experiment support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in rescue experiment support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating rescue experiment support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Safety Biomarker Validity
The degree to which safety biomarker supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000054
BEMO:2000054
61
a6b6c6cf8ada87a1112583df9bc78595f4f0975ec0e8bbfc01ca6f00c2c0adf1
10
Assesses safety biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in safety biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating safety biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Collection Procedure Consistency
The degree of agreement in collection procedure across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000064
BEMO:2000064
71
c748a49d52923f74e8e31a0f69a4e91ee50970289d8b43c903a638e0573cc828
8
Assesses collection procedure consistency using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in collection procedure consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating collection procedure consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Materials-and-Reagents Reporting Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Patient Flow Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of patient flow integrity.
bemo
BEMO:2000126
BEMO:2000126
133
67544e05c38d184f5e9fa9830136b1408e7835e115b1fa9c54147ad65e14b99c
8
Assesses patient flow integrity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in patient flow integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating patient flow integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 426
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Quality-Control Reporting Completeness; Null-Result Interpretability; Deviations-from-Protocol Transparency", "Common Misinterpretations": "Treating negative-result reporting as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Negative-Result Reporting", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of negative-result reporting.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses negative-result reporting using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in negative-result reporting can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
426
f6803e55dc7d7550d62e5bed1f8773c7da7d0266b37145da5354a96ceaf769c2
Baseline Comparability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of baseline comparability.
bemo
BEMO:2000466
BEMO:2000466
473
93fad13257c869638e96dc16ab4687d4b955a0d40d8d10736f49f12cebec65eb
8
Assesses baseline comparability using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in baseline comparability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating baseline comparability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Noninferiority Margin Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Therapeutic Window Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Practical Identifiability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Cross-Platform Genomic Concordance
The degree of agreement in cross-platform genomic across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000243
BEMO:2000243
250
91130249c6d3c5254de370a79251a5dce82ecc1afbe7b1276235f59ee93b7b5a
8
Assesses cross-platform genomic concordance using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in cross-platform genomic concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cross-platform genomic concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Result Reproducibility
The degree to which result reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000405
BEMO:2000405
412
60121a2998efa943018d3ff7dd4130d4332ee279da0e8111b650ea175e9dacce
10
Assesses result reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in result reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating result reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Protocol Deviation Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Parameter Identifiability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Deviations-from-Protocol Transparency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Code Availability
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Calibration of Statistical Predictions
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of calibration of statistical predictions.
bemo
BEMO:2000433
BEMO:2000433
440
1d5cbda1b0dc5b30aa8339e6219d08168170bd684f74bf0679b26dd4ece9ead4
8
Assesses calibration of statistical predictions using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in calibration of statistical predictions can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating calibration of statistical predictions as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 325
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Relation Evidence Strength; Ontology Evidence-Code Strength; Model–Experiment Concordance", "Common Misinterpretations": "Treating ontology annotation completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Ontology Annotation Completeness", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The extent to which all scientifically necessary components of ontology annotation are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses ontology annotation completeness using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in ontology annotation completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
325
b8714924720c209e415f9bf2c9aaf707f6bf0edbbbf6ca8da11d2546f9da563d
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 272
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Normalization Adequacy; Alternative Splicing Validation; Single-Cell Doublet Burden", "Common Misinterpretations": "Treating transcript quantification reliability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Transcript Quantification Reliability", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of transcript quantification reliability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses transcript quantification reliability using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in transcript quantification reliability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
272
45c762be688f8b3f57f3b563b58fc816d5802e731de9aee9acb732eea429ea1a
Data-Sharing Transparency
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of data-sharing transparency.
bemo
BEMO:2000410
BEMO:2000410
417
24a8774755f1b71359ef7a35f4a7813a446670f2da0ccc38d22836e1a84f1453
8
Assesses data-sharing transparency using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in data-sharing transparency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating data-sharing transparency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Discrimination of Statistical Predictions
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Performance Bias Risk
The probability or degree that performance bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000481
BEMO:2000481
488
67d1ca80d7087e3cfba7c1c9f64d45d3f4094887ac27aef09ef61bfb6d49d732
10
Assesses performance bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in performance bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating performance bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Computation profile for Data Provenance Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Age Appropriateness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Repeatability
The degree to which repeatability yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000297
BEMO:2000297
304
289ee6101d32dc64bb4644380815480f0d1de4b3d0512c355bf6d1bb8e30c7f2
8
Assesses repeatability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in repeatability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating repeatability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Cross-Reactivity
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 262
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Population Stratification Control; Differential Expression Robustness; Normalization Adequacy", "Common Misinterpretations": "Treating relatedness control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Relatedness Control", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of relatedness control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses relatedness control using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in relatedness control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
262
c66c714b8c2b1fb9da08eabc61ab1b346b5733232f7901baf90ed1206b3491ba
Computation profile for Microbial Contamination
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Reportable Range
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of reportable range.
bemo
BEMO:2000298
BEMO:2000298
305
1958a10669f076f26b0e434840fd948f2a615d7e3a768038db89230b0e148857
8
Assesses reportable range using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in reportable range can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reportable range as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 290
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Matrix Effect; Cross-Reactivity; Carryover", "Common Misinterpretations": "Treating interference susceptibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Interference Susceptibility", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of interference susceptibility.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses interference susceptibility using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in interference susceptibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
290
af7341a32f2656e5638dc5bce7ac126974e88c78554d2916a696247ebbe11d2b
Computation profile for Prediction Interval Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
scientific importance score
Source-provided scientific importance score.
BEMO:3100024
Linearity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of linearity.
bemo
BEMO:2000287
BEMO:2000287
294
511e3fb111714ad52a916859d22dfd9a9aa36539d615e5b295fdf52500ac32e5
8
Assesses linearity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in linearity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating linearity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Carryover
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of carryover.
bemo
BEMO:2000275
BEMO:2000275
282
0e987b7e93f5f6b586140ce3b08eb3d87acd863bc4cb825ee0d1e4b7b1bb3994
8
Assesses carryover using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in carryover can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating carryover as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Calibration Traceability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Humane Endpoint Appropriateness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Endpoint Reliability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of endpoint reliability.
bemo
BEMO:2000042
BEMO:2000042
49
e147bf4116de09d3c1bd6b5fb8081223589d31f54f697b6e869790830619757d
8
Assesses endpoint reliability using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in endpoint reliability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating endpoint reliability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Attrition Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Time-Dependent Discrimination
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of time-dependent discrimination.
bemo
BEMO:2000137
BEMO:2000137
144
5232bfe9f1dc1ddd88662ddb65a550d0d1fc5338bb4a530ed1072ed6078116bf
8
Assesses time-dependent discrimination using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in time-dependent discrimination can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating time-dependent discrimination as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Independent Replication Strength
The magnitude and credibility of independent evidence supporting independent replication.
bemo
BEMO:2000394
BEMO:2000394
401
3837461a84fca58515b90e02475b1cf195c845c2cf4704aa9d0c7c2b412d9406
10
Assesses independent replication strength using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in independent replication strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating independent replication strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Protocol Fidelity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of protocol fidelity.
bemo
BEMO:2000484
BEMO:2000484
491
ae404e8a9df9a665347ede4db7ae7370bdc6c77c15da3474e23c516ad81dd25e
8
Assesses protocol fidelity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in protocol fidelity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating protocol fidelity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Toxicological Mode-of-Action Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 322
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Network Reconstruction Robustness; Network Node Confidence; Module Stability", "Common Misinterpretations": "Treating network edge confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Network Edge Confidence", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The justified degree of certainty assigned to network edge given the quantity, quality, consistency, and limitations of supporting evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses network edge confidence using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in network edge confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
322
6118b228a2af4e149e82150fb026ef08cdd6d05b9748a0128d6e8c14f147784c
Computation profile for Comparative Test Accuracy
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Control Group Appropriateness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Phenotypic Specificity for Variant
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of phenotypic specificity for variant.
bemo
BEMO:2000229
BEMO:2000229
236
ae0b81578c2487ac58cc25c0925e5e5b5580990b17b1391fc512e266a26a2788
9
Assesses phenotypic specificity for variant using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in phenotypic specificity for variant can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating phenotypic specificity for variant as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 287
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Reportable Range; Recovery; Dilution Integrity", "Common Misinterpretations": "Treating dynamic range as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Dynamic Range", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of dynamic range.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Method- and analyte-specific physical units", "What It Measures": "Assesses dynamic range using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in dynamic range can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
287
41309e5f36c397d08d76621c273137352b099c93484f2a2eca025bcf85ccffdc
Batch-Effect Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of batch-effect control.
bemo
BEMO:2000239
BEMO:2000239
246
0751d3491af658db1257b3b04a366a751c7fd0fdfb5e615ebbae3dab6970c7c9
8
Assesses batch-effect control using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in batch-effect control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating batch-effect control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Biological Gradient
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biological gradient.
bemo
BEMO:2000001
BEMO:2000001
8
6471dcf670c9dd05424f4606389decf6b51812d07c0a0680842d77bda957b0f3
8
Assesses biological gradient using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in biological gradient can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biological gradient as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 32
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Epistasis Support; On-Target Specificity; Off-Target Liability Evidence", "Common Misinterpretations": "Treating target engagement evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Target Engagement Evidence", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of target engagement evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses target engagement evidence using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in target engagement evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
32
a51c69d24f2bf9ef711f1d3b54d2086032245754bf1992c1b5999a88e3ca185f
Computation profile for Evidence Coherence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Partial Area Under the Receiver Operating Characteristic Curve
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of partial area under the receiver operating characteristic curve.
bemo
BEMO:2000125
BEMO:2000125
132
79f6b607e0d7504f46e1d3842fa8966ae929db7058236a6b4ed553277ad444b5
8
Assesses partial area under the receiver operating characteristic curve using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in partial area under the receiver operating characteristic curve can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating partial area under the receiver operating characteristic curve as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Incremental Diagnostic Value
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Funding-Source Transparency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Participant Flow Completeness
The extent to which all scientifically necessary components of participant flow are present, documented, and evaluable.
bemo
BEMO:2000422
BEMO:2000422
429
c3caaa14ecea68f7075a25ef235cdae89d1a9b9b86ed2b5a30dab5cb5b8bf20d
9
Assesses participant flow completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in participant flow completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating participant flow completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Carryover
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 39
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Biomarker Responsiveness; Known-Groups Validity; Convergent Validity", "Common Misinterpretations": "Treating biomarker sensitivity to change as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Biomarker Sensitivity to Change", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biomarker sensitivity to change.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses biomarker sensitivity to change using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biomarker sensitivity to change can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
39
80beed232e03f65da1843b394120d8baef7fb4428401374caac926260837e2b2
Selectivity Profile
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of selectivity profile.
bemo
BEMO:2000353
BEMO:2000353
360
2957579a3d734fc808c8a89eb160fb182b5f5a0f68ba40c4eecc1921dea93153
8
Assesses selectivity profile using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in selectivity profile can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating selectivity profile as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Time–Concentration Profile Adequacy
The extent to which time–concentration profile is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000356
BEMO:2000356
363
8bfcea065494f7118cb14f95a9459e3f35fa404437872877955c4fdff069147a
8
Assesses time–concentration profile adequacy using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in time–concentration profile adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating time–concentration profile adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 14
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Biological Gradient; Network Context Support; Systems-Level Emergence Support", "Common Misinterpretations": "Treating homeostatic compensation assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Homeostatic Compensation Assessment", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of homeostatic compensation assessment.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses homeostatic compensation assessment using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in homeostatic compensation assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
14
61aaa3d7ba71760f01fcbbb79304031b61f416940944ef459f48b6d07dc0628d
Therapeutic Window Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of therapeutic window evidence.
bemo
BEMO:2000355
BEMO:2000355
362
34f5d4730ddbb7e269fc6713d96c49c0fba2d2f738ddbfb47d339a6ed5f8cf6e
8
Assesses therapeutic window evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in therapeutic window evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating therapeutic window evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Data Availability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of data availability.
bemo
BEMO:2000390
BEMO:2000390
397
89d3d59a31fe591c17a3c61f83a55d18326e4b311b8033047e95bfb2a93715fa
8
Assesses data availability using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in data availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating data availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Population Stratification Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 166
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Information Size Adequacy", "Common Misinterpretations": "Treating multiplicity-adjusted credibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Multiplicity-Adjusted Credibility", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of multiplicity-adjusted credibility.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses multiplicity-adjusted credibility using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in multiplicity-adjusted credibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
166
40900cc3965adf1a4a61ca1880158e0ed27ede5d1a6163f63eb4f6d38b99a3b8
Computation profile for Null-Result Interpretability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Missing-Data Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Potency Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Reference Bias
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of reference bias.
bemo
BEMO:2000254
BEMO:2000254
261
7f170b7ac0b632bb15937074f6f54297bd1374274ba7dd6a71c308a1270c144b
10
Assesses reference bias using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in reference bias can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reference bias as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Study Heterogeneity
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Sample Identity Concordance
The degree of agreement in sample identity across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000256
BEMO:2000256
263
cf22f061ff0a6f0403c89f061c9c86c0fdc0e5186e885958801878829549f2cd
8
Assesses sample identity concordance using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in sample identity concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating sample identity concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Prospective Registration
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prospective registration.
bemo
BEMO:2000425
BEMO:2000425
432
32ae2352b0f17339d30893a08c10c681c7fa3992e6b60a3343dae6e95aa8dda5
8
Assesses prospective registration using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in prospective registration can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating prospective registration as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Interlaboratory Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 38
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Biomarker Reliability; Biomarker Sensitivity to Change; Known-Groups Validity", "Common Misinterpretations": "Treating biomarker responsiveness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Biomarker Responsiveness", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biomarker responsiveness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses biomarker responsiveness using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biomarker responsiveness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
38
19a54854b5dfb7168eddad304da95ddbab2e1ad34992146fcedfd650dfe09548
Index-Test Blinding
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of index-test blinding.
bemo
BEMO:2000121
BEMO:2000121
128
3bc623a30f8cb3956f6d71196888c1021137e72c87604513c9fc1e352f98bb4c
8
Assesses index-test blinding using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in index-test blinding can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating index-test blinding as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Evidence Strength
The magnitude and credibility of independent evidence supporting evidence.
bemo
BEMO:2000154
BEMO:2000154
161
d0050ec2d0692a0a792ec941fce4c076cf8613142effd9ef06a03cab99fa82fe
8
Assesses evidence strength using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Sample Stability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of sample stability.
bemo
BEMO:2000300
BEMO:2000300
307
42f10127b0bb2f21e12bfaaba11ee2a6c8335cd39935669d6733741eadfc7015
9
Assesses sample stability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in sample stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating sample stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Temporal Precedence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 293
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Limit of Detection; Linearity; Analytical Measurement Range", "Common Misinterpretations": "Treating limit of quantification as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Limit of Quantification", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of limit of quantification.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Method- and analyte-specific physical units", "What It Measures": "Assesses limit of quantification using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in limit of quantification can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
293
e110a21f976b308d8a01ff42efac55157e3af426e1182c74cf700e2fb4ce3595
Domain Protocol Required
Controlled BEMO ComputationReadinessStatus value: DomainProtocolRequired.
BEMO:4000013
DomainProtocolRequired
Temporal Precedence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of temporal precedence.
bemo
BEMO:2000488
BEMO:2000488
495
ec85a4abfc9e480fea03dfff5177f17a293899d488f06a679715de258c8591b9
8
Assesses temporal precedence using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in temporal precedence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating temporal precedence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Dose Proportionality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of dose proportionality.
bemo
BEMO:2000334
BEMO:2000334
341
6d7bafb114f62ee059333139662f0dfa2a0fb0aaab829ac8fcdb03ed2fbd0708
8
Assesses dose proportionality using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in dose proportionality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating dose proportionality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Direct Replication Success
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of direct replication success.
bemo
BEMO:2000392
BEMO:2000392
399
e6b0ec1b803bb2c7cb8becab175dbb88ab91c9d8ccd3f713a814c836846ea547
10
Assesses direct replication success using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in direct replication success can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating direct replication success as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Setting Applicability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of setting applicability.
bemo
BEMO:2000206
BEMO:2000206
213
323f595d53815c626b380456c55c73f7ff5e2e10beb1e0ffc44edb8067c37e86
8
Assesses setting applicability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in setting applicability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating setting applicability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Contamination Risk
The probability or degree that contamination introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000472
BEMO:2000472
479
eed78db6e72161c41f0a51c8bc55e5d20b0d6675373ef1a7bf65b4c5b0dc75c0
8
Assesses contamination risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in contamination risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating contamination risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Fragmentation Spectrum Quality
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 326
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Ontology Annotation Completeness; Model–Experiment Concordance; Parameter Identifiability", "Common Misinterpretations": "Treating ontology evidence-code strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Ontology Evidence-Code Strength", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The magnitude and credibility of independent evidence supporting ontology evidence-code.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses ontology evidence-code strength using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in ontology evidence-code strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
326
98cc3b5d977b0d0bd7a309c477d2653f05eeb7554078021a2d0b2bf21036d891
Computation profile for Parameter Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 118
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Prognostic Calibration; Calibration-in-the-Large; Observed-to-Expected Ratio", "Common Misinterpretations": "Treating calibration slope as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Calibration Slope", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of calibration slope.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses calibration slope using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in calibration slope can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
118
44bfcfec74d2b98de3507998b5e63f6506912c93210fea4df7a678aed4417a75
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 305
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Analytical Measurement Range; Dynamic Range; Recovery", "Common Misinterpretations": "Treating reportable range as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Reportable Range", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of reportable range.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Method- and analyte-specific physical units", "What It Measures": "Assesses reportable range using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reportable range can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
305
1958a10669f076f26b0e434840fd948f2a615d7e3a768038db89230b0e148857
Counterevidence Strength
The magnitude and credibility of independent evidence supporting counterevidence.
bemo
BEMO:2000140
BEMO:2000140
147
d4c1d3bbceb7e86e058c678a7bbc386dd5980f93dcbc9a60ebc6c6d39231fcdb
8
Assesses counterevidence strength using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in counterevidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating counterevidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Emergent-Property Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Developing; BEMO computation profile requires independent validation.
Diagnostic Odds Ratio
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of diagnostic odds ratio.
bemo
BEMO:2000115
BEMO:2000115
122
a903f03627bce43e756f3c1d10c27c552a53bce0346a40196ea1cd72afe63de0
8
Assesses diagnostic odds ratio using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in diagnostic odds ratio can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.
Treating diagnostic odds ratio as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Matched-Sample Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of matched-sample integrity.
bemo
BEMO:2000071
BEMO:2000071
78
daf35eaf7e6632e499a494b9ba629171cde4aa96d4d637ad7669da2f69668fd5
8
Assesses matched-sample integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in matched-sample integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating matched-sample integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Genotoxicity Evidence Strength
The magnitude and credibility of independent evidence supporting genotoxicity evidence.
bemo
BEMO:2000338
BEMO:2000338
345
e40592e267890e51ad35407c7dc2af7775fdac9174afe2f08486d28a5e2901b8
8
Assesses genotoxicity evidence strength using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in genotoxicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating genotoxicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 114
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Causal Contrast Clarity; Instrument Validity; Negative-Control Validation", "Common Misinterpretations": "Treating target trial emulation fidelity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Target Trial Emulation Fidelity", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of target trial emulation fidelity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses target trial emulation fidelity using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in target trial emulation fidelity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
114
7a907a8dc4fe2c40df0aa3c333c85c8e7231c1e33d26321171335bcd3464b6a6
Computation profile for Phenotypic Specificity for Variant
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Mechanistic Specificity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of mechanistic specificity.
bemo
BEMO:2000013
BEMO:2000013
20
29074cf2dac765628d0a639b922f2e3313846aa414af9c4ef760ea93c86819ca
9
Assesses mechanistic specificity using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in mechanistic specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating mechanistic specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Network Reconstruction Robustness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 451
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Outlier Influence Robustness; Nonlinearity Assessment; Interaction Assessment Adequacy", "Common Misinterpretations": "Treating influential observation sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Influential Observation Sensitivity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of influential observation sensitivity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses influential observation sensitivity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in influential observation sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
451
0202d30b443e9ba2dd8ad4597c8e7b9fcc240af0f60b19db4a115b1d2866e858
Computation profile for Eligibility Criteria Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Mechanistic Specificity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Anatomical Site Fidelity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Deviations-from-Protocol Transparency
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of deviations-from-protocol transparency.
bemo
BEMO:2000411
BEMO:2000411
418
6d08df9b94938aa4356ba3bbe963ef6fed2e0e499e09e30936b635d55a24d3c8
8
Assesses deviations-from-protocol transparency using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in deviations-from-protocol transparency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating deviations-from-protocol transparency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Biomarker Clinical Relevance
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biomarker clinical relevance.
bemo
BEMO:2000028
BEMO:2000028
35
70bbbca5f05350dbfed193659c18ff0eef66a6552528618a931911c0bbfe7878
10
Assesses biomarker clinical relevance using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in biomarker clinical relevance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biomarker clinical relevance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Commutability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of commutability.
bemo
BEMO:2000276
BEMO:2000276
283
0d8876821c8ce33e4a6f72fcdd334c759d21fb2784bd7f9f779179758d7849c8
8
Assesses commutability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in commutability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating commutability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 164
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Cumulative Evidence Stability; Multiplicity-Adjusted Credibility", "Common Misinterpretations": "Treating information size adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Information Size Adequacy", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The extent to which information size is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses information size adequacy using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in information size adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
164
1349313eaf86ad90454f8dbf26aa9cb5389132fc9ac2e44b703024851df5136c
Computation profile for Experimental Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Flux-Balance Consistency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
directionality
A controlled concept describing how greater or lower values should be interpreted.
bemo
BEMO:0000404
Candidate
Computation profile for Limit of Quantification
0.1.0
Apply a validated analyte- and method-specific measurement procedure with calibration and quality control.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","unitRef":"REQUIRED"}
specimen_or_material; measurement_procedure; calibration_reference; quality_control_results; unit
replicate_measurements; environmental_conditions; instrument_version
xsd:decimal
Method- and analyte-specific physical units
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Hotspot/Functional-Domain Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Mixture Interaction Assessment
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Pharmacological Target Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Cross-Species Biological Concordance
The degree of agreement in cross-species biological across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000004
BEMO:2000004
11
6f883b8ebe7aa3d7407688c8b2390b245a894c21f5c21ffef99fe634674c5fd3
8
Assesses cross-species biological concordance using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in cross-species biological concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cross-species biological concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Duplicate Read Burden
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 198
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Disease-Severity Generalizability; Context Sensitivity; Ecological Validity", "Common Misinterpretations": "Treating care-pathway independence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Care-Pathway Independence", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of care-pathway independence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses care-pathway independence using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in care-pathway independence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
198
e39c9004563f895f591d754f0c2c01d5cfe45a0e26d44584ecb91c07d72c494a
Gain-of-Function Validation
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of gain-of-function validation.
bemo
BEMO:2000006
BEMO:2000006
13
5c02cfac370a9d2c1514479546d6e291e60bffe1758412ceec02bbf8edaab2dc
8
Assesses gain-of-function validation using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in gain-of-function validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating gain-of-function validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Matrix Effect
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Epistasis Support
The magnitude and credibility of independent evidence supporting epistasis.
bemo
BEMO:2000005
BEMO:2000005
12
cd57f610728234561644257a8387987f0aa8934688593c5fbd817b3779032054
8
Assesses epistasis support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in epistasis support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating epistasis support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Sampling Frame Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 250
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Spatial Transcriptomic Registration Accuracy", "Common Misinterpretations": "Treating cross-platform genomic concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Cross-Platform Genomic Concordance", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "The degree of agreement in cross-platform genomic across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses cross-platform genomic concordance using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cross-platform genomic concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
250
91130249c6d3c5254de370a79251a5dce82ecc1afbe7b1276235f59ee93b7b5a
Multi-omics and Systems Biology metric
Category of biomedical evidence metrics concerned with multi-omics and systems biology.
bemo
BEMO:1100012
Candidate
Computation profile for Sex Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Differential Expression Robustness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Developmental Toxicity Evidence Strength
The magnitude and credibility of independent evidence supporting developmental toxicity evidence.
bemo
BEMO:2000333
BEMO:2000333
340
f5636ed02b8746eb48d98671b027c695a8940399bff60ebeae8ca955d99bae6a
8
Assesses developmental toxicity evidence strength using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in developmental toxicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating developmental toxicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Type I Error Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of type i error control.
bemo
BEMO:2000460
BEMO:2000460
467
a3eba05a4613d95dc796e3df689756346c4f78ac2fda1dcb4e8cf4b813a22638
8
Assesses type i error control using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in type i error control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating type i error control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Benchmark Dose Reliability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of benchmark dose reliability.
bemo
BEMO:2000330
BEMO:2000330
337
43d4a4ce7127bc52167a879f611c41413072f25e775dbe7c9446cf4028783b10
8
Assesses benchmark dose reliability using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in benchmark dose reliability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating benchmark dose reliability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Allocation Concealment
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of allocation concealment.
bemo
BEMO:2000464
BEMO:2000464
471
3367e7e58461dcb460260b263a504d906f8c26af8608d0cf2767790d3843a2e7
8
Assesses allocation concealment using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in allocation concealment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating allocation concealment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Multiplicity Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of multiplicity control.
bemo
BEMO:2000451
BEMO:2000451
458
282fd29ea793d9f6346d4b34598e7ca30406dad2a72b3f3167d354ca92803ec0
8
Assesses multiplicity control using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in multiplicity control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating multiplicity control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 268
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Single-Cell Viability; Cell-Type Annotation Confidence; Spatial Transcriptomic Registration Accuracy", "Common Misinterpretations": "Treating single-cell feature detection rate as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Single-Cell Feature Detection Rate", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of single-cell feature detection rate.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses single-cell feature detection rate using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in single-cell feature detection rate can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
268
1ce87bfa0e14b3c49cdf26e92f3b39ba4ce32c6192d394a99b4f3b80c80031b2
Computation profile for Publication Bias Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Preservation Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Gain-of-Function Validation
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Recruitment Reporting Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 17
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Mechanistic Completeness; Mechanistic Specificity; Mechanistic Causality Strength", "Common Misinterpretations": "Treating mechanistic coherence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Mechanistic Coherence", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of mechanistic coherence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses mechanistic coherence using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in mechanistic coherence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
17
dc1b3b375698d0aec42693d5a0b2f87405e7e7bf90759105797ad714ab4e9ab2
Endpoint Validity
The degree to which endpoint supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000044
BEMO:2000044
51
f79c486fa1f56b17619593da318c44a3704cbd1294be41ae30a2b38ceb9410d1
10
Assesses endpoint validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in endpoint validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating endpoint validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Bayes Factor Evidence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Vehicle-Control Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 194
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Animal Model Predictive Validity; Sex as a Biological Variable Adequacy; Age Appropriateness", "Common Misinterpretations": "Treating species appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Species Appropriateness", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The extent to which species is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses species appropriateness using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in species appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
194
93f7de14ca784e8a47e3c6ab351c1e75c21a4c3e26305cfe63e0d4fef2122575
Calibration-in-the-Large
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of calibration-in-the-large.
bemo
BEMO:2000112
BEMO:2000112
119
4a117ceb82993be69cf387d42f6626e25d4ff682e98eb458d01d55a60ca6f29d
8
Assesses calibration-in-the-large using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in calibration-in-the-large can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating calibration-in-the-large as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Dechallenge–Rechallenge Support
The magnitude and credibility of independent evidence supporting dechallenge–rechallenge.
bemo
BEMO:2000091
BEMO:2000091
98
9e166f37b12b35c895b18fc7e56e04526e22004c0b70dee562c1cf80c5d17e69
8
Assesses dechallenge–rechallenge support using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in dechallenge–rechallenge support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating dechallenge–rechallenge support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Homeostatic Compensation Assessment
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Species Appropriateness
The extent to which species is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000187
BEMO:2000187
194
93f7de14ca784e8a47e3c6ab351c1e75c21a4c3e26305cfe63e0d4fef2122575
8
Assesses species appropriateness using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in species appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating species appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 275
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Linearity; Reportable Range; Dynamic Range", "Common Misinterpretations": "Treating analytical measurement range as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Analytical Measurement Range", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of analytical measurement range.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Method- and analyte-specific physical units", "What It Measures": "Assesses analytical measurement range using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in analytical measurement range can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
275
40562e2bdf0f0c373312aa0a75941fd6a83f056ba5201b300a46242b3f5457df
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 107
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Instrument Validity; Dose–Response Support; Dechallenge–Rechallenge Support", "Common Misinterpretations": "Treating negative-control validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Negative-Control Validation", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of negative-control validation.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses negative-control validation using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in negative-control validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
107
63bedc3fc09ef48160a8d522d76618e50e47197d3d502ae109f1015d0fbcb67e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 313
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Practical Identifiability; Steady-State Validity; Flux-Balance Consistency", "Common Misinterpretations": "Treating dynamical stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Dynamical Stability", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of dynamical stability.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses dynamical stability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in dynamical stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
313
d58852cf510c70b8fcfa12f792791c36728adc16d59c5b9bb4b7a9b129c39f52
Computation profile for Mapping Quality
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Analytical Specificity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of analytical specificity.
bemo
BEMO:2000271
BEMO:2000271
278
f6458dae6787eecc6dbfe775e26236c89d02f80c2f92e3575c7f9ddfa0e84b0e
10
Assesses analytical specificity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in analytical specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating analytical specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Method Comparison Agreement
The degree of agreement in method comparison across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000290
BEMO:2000290
297
8f2dfbddd616818e03a89990f48f31b1db69460f1919bb65cc71ee0d68b7b4f6
8
Assesses method comparison agreement using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in method comparison agreement can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating method comparison agreement as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Clinical Relevance of Effect
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of clinical relevance of effect.
bemo
BEMO:2000434
BEMO:2000434
441
7e04bf59de003d9389f3d5e5a39c09dbec7b25b54fdd665e24d339ccf148fb17
10
Assesses clinical relevance of effect using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in clinical relevance of effect can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating clinical relevance of effect as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 78
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Chain-of-Custody Integrity", "Common Misinterpretations": "Treating matched-sample integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Matched-Sample Integrity", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of matched-sample integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses matched-sample integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in matched-sample integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
78
daf35eaf7e6632e499a494b9ba629171cde4aa96d4d637ad7669da2f69668fd5
Protocol Availability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of protocol availability.
bemo
BEMO:2000426
BEMO:2000426
433
6e6dfc0138818017108063ab02fea9da669c25ffc36ec988048eb058c538f84a
8
Assesses protocol availability using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in protocol availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating protocol availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Locus Heterogeneity Assessment
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Cross-Omics Integration Coherence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Developing; BEMO computation profile requires independent validation.
Analytical Precision
The closeness of repeated estimates or measurements and the narrowness of uncertainty around analytical.
bemo
BEMO:2000269
BEMO:2000269
276
7c4987f5f8c3c9cca41bb28896f2392e5c664fc505d3f5db2c63ad7c10e1d04d
10
Assesses analytical precision using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in analytical precision can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating analytical precision as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Discriminant Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 260
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Mapping Quality; Duplicate Read Burden; Library Complexity", "Common Misinterpretations": "Treating read quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Read Quality", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of read quality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses read quality using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in read quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
260
e9483af6db486496b6f4b0c18d333bb38a6c500a6b80dd23beefe82c05de4acc
has applicability profile
Connects a metric to an applicability profile.
BEMO:3000013
Transportability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of transportability.
bemo
BEMO:2000210
BEMO:2000210
217
ef0b46b07942a1780f9e2e890b09b7edc3579f016762848aa46c4d956da3075d
8
Assesses transportability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in transportability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating transportability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Metabolite Annotation Level
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of metabolite annotation level.
bemo
BEMO:2000370
BEMO:2000370
377
883e5fdd2ff9e198b3ddcbafc3e6478eed144bbcc7f46c67f19ec93850e8ec4c
8
Assesses metabolite annotation level using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in metabolite annotation level can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating metabolite annotation level as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
metric output specification
An information content entity specifying the datatype, scale, unit, range, and interpretation of a metric output.
bemo
BEMO:0000102
Candidate
Instrument Drift
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of instrument drift.
bemo
BEMO:2000282
BEMO:2000282
289
39eec82dab659b7826d3bea8f3227e8afeb55e0f15214b02955200cb8a59a8ba
8
Assesses instrument drift using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in instrument drift can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating instrument drift as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for On-Target Specificity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Protein Identification Confidence
The justified degree of certainty assigned to protein identification given the quantity, quality, consistency, and limitations of supporting evidence.
bemo
BEMO:2000377
BEMO:2000377
384
86cf8ad175ebd274f625e908935f946750c2c41a33f40e7a8dc33224d8db7d28
10
Assesses protein identification confidence using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in protein identification confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating protein identification confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Known-Groups Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 479
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Protocol Deviation Risk; Co-intervention Bias Risk; Carryover Effect Risk", "Common Misinterpretations": "Treating contamination risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Contamination Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that contamination introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses contamination risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in contamination risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
479
eed78db6e72161c41f0a51c8bc55e5d20b0d6675373ef1a7bf65b4c5b0dc75c0
source workbook
Name of the source workbook.
BEMO:3200014
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 75
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Necrosis Burden; Lipemia Burden; Icterus Interference", "Common Misinterpretations": "Treating hemolysis burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Hemolysis Burden", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of hemolysis burden.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses hemolysis burden using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in hemolysis burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
75
1b0946c6b109834d9e6fa6dc5ae7dde1ce95bc65f999a58a3c5a88bc66c6a47c
Interaction Assessment Adequacy
The extent to which interaction assessment is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000445
BEMO:2000445
452
7942499ba472159741e9410722d92887d17131b3d661f12b98b8ec8f2d490791
8
Assesses interaction assessment adequacy using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in interaction assessment adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating interaction assessment adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Estimate Precision
The closeness of repeated estimates or measurements and the narrowness of uncertainty around estimate.
bemo
BEMO:2000441
BEMO:2000441
448
9458d6e4371d11a80963339d3a9a6c5dcc187a36080cf18d964ecb5827c41820
9
Assesses estimate precision using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in estimate precision can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating estimate precision as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Computation profile for Necrosis Burden
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Evidence Coverage
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 226
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Expressivity Consistency; Phenocopy Risk; Locus Heterogeneity Assessment", "Common Misinterpretations": "Treating founder-effect assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Founder-Effect Assessment", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of founder-effect assessment.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses founder-effect assessment using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in founder-effect assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
226
6530e53b02059e4bfd6465ae64ac385198a560d8a80df4ba701894a8c5cb0975
Computation profile for Orthogonal Validation
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Null-Result Interpretability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of null-result interpretability.
bemo
BEMO:2000420
BEMO:2000420
427
0e7be34e4db71c86201d878decffa3c8da8ca62a1c8481dc3779d3c4caa454b1
8
Assesses null-result interpretability using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in null-result interpretability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating null-result interpretability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 396
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Direct Replication Success; Analytical Reproducibility; Computational Reproducibility", "Common Misinterpretations": "Treating conceptual replication success as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Conceptual Replication Success", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of conceptual replication success.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses conceptual replication success using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in conceptual replication success can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
396
45fcdc0cd9da1d73796f886d24683b4f40ead7f7a14a06e84d39631b66a910ad
Computation profile for Steady-State Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 461
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Imputation Validity; Influential Observation Sensitivity; Nonlinearity Assessment", "Common Misinterpretations": "Treating outlier influence robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Outlier Influence Robustness", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of outlier influence robustness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses outlier influence robustness using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in outlier influence robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
461
3a52b33f3faa79b95100ff084b9c7ca54158523c656fcaf5c298f073960428c9
Peptide Identification Confidence
The justified degree of certainty assigned to peptide identification given the quantity, quality, consistency, and limitations of supporting evidence.
bemo
BEMO:2000374
BEMO:2000374
381
7f78eede8aa05781dde3f12d7ea8111e1c0b552f5b722c402b0de5323e65b0ea
10
Assesses peptide identification confidence using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in peptide identification confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating peptide identification confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Metadata Completeness
The extent to which all scientifically necessary components of metadata are present, documented, and evaluable.
bemo
BEMO:2000417
BEMO:2000417
424
cfce2b3ca58f108e1dc5461ad1ef2f5a84b55171c75536d306bf4e1e2b3bc58d
9
Assesses metadata completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in metadata completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating metadata completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Subgroup Consistency
The degree of agreement in subgroup across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000208
BEMO:2000208
215
26cabc1a909ce4ef7d5d54dc609f85f9fea3829985a8f8573b3da07a3e225f5b
8
Assesses subgroup consistency using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in subgroup consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating subgroup consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Icterus Interference
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Prediction Interval Adequacy
The extent to which prediction interval is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000161
BEMO:2000161
168
0acc1464b5d065a6f38744191dff4d933036cf918aeac3f16274b9cdb9dd1059
8
Assesses prediction interval adequacy using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in prediction interval adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating prediction interval adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Reanalysis Concordance
The degree of agreement in reanalysis across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000403
BEMO:2000403
410
07f0b47b85e307e7887d3c938a89cbcad651d37d8bfa888ea4ce710d25410a29
8
Assesses reanalysis concordance using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in reanalysis concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating reanalysis concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
biomedical evidence metric
An information content entity that specifies a scientifically interpretable quantity, category, rubric, or assessment construct used to evaluate biomedical evidence.
bemo
BEMO:0000001
Candidate
Method Specific Protocol Required
Controlled BEMO FormulaStatus value: MethodSpecificProtocolRequired.
BEMO:4000020
MethodSpecificProtocolRequired
Outcome Definition Completeness
The extent to which all scientifically necessary components of outcome definition are present, documented, and evaluable.
bemo
BEMO:2000421
BEMO:2000421
428
129e8fff9350e8d3fb88d0a300e3816f498afcc03cf17b3f6037da094441e4fe
9
Assesses outcome definition completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in outcome definition completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating outcome definition completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Outcome Relevance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Context-of-Use Validity
The degree to which context-of-use supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000037
BEMO:2000037
44
170c4bee423f9901b11abd5c6c7203f1253da40dbd46a33c459bac07471058e8
10
Assesses context-of-use validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in context-of-use validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating context-of-use validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Quantitative Bias Analysis Robustness
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Fixation Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Attrition Bias Risk
The probability or degree that attrition bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000465
BEMO:2000465
472
0783cbbb1a7caf7ef448d6830ae564e73010beb87e2e08710ddb494b6b76ccb3
10
Assesses attrition bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in attrition bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating attrition bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Drug–Drug Interaction Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of drug–drug interaction evidence.
bemo
BEMO:2000335
BEMO:2000335
342
3a82aa8a1751fa8372cd9735fd975781682a359be68aacfb1d0d78e896d96ce5
8
Assesses drug–drug interaction evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in drug–drug interaction evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating drug–drug interaction evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Equivalence Margin Validity
The degree to which equivalence margin supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000440
BEMO:2000440
447
226b21fe561d36c12bfde0bcc47bb1ad6cb02760533a4b768747d6a0ab92184b
10
Assesses equivalence margin validity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in equivalence margin validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating equivalence margin validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 230
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Loss-of-Function Mechanism Validity; Null-Variant Quality; Splicing Evidence Strength", "Common Misinterpretations": "Treating hotspot/functional-domain evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Hotspot/Functional-Domain Evidence", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of hotspot/functional-domain evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses hotspot/functional-domain evidence using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in hotspot/functional-domain evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
230
d2e72ca269dc077aef430f8efab1ad3e7adf8db5e6aa339468796f69dfaaee12
Computation profile for Negative-Control Performance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Efficacy Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Endpoint Responsiveness
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Post-Translational Modification Localization Confidence
The justified degree of certainty assigned to post-translational modification localization given the quantity, quality, consistency, and limitations of supporting evidence.
bemo
BEMO:2000376
BEMO:2000376
383
6761ab89a934b85357a2d58a5e53503065037a8cd658104babdb5536291a3653
10
Assesses post-translational modification localization confidence using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in post-translational modification localization confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating post-translational modification localization confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Variant Call Quality
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Intermediate Precision
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Evidence Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Read Quality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of read quality.
bemo
BEMO:2000253
BEMO:2000253
260
e9483af6db486496b6f4b0c18d333bb38a6c500a6b80dd23beefe82c05de4acc
8
Assesses read quality using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in read quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.
Treating read quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Evidence Directness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence directness.
bemo
BEMO:2000148
BEMO:2000148
155
880de5f192c9dc60c14576e428f1c46fbdcda3180afe950bcde297920e65c6df
8
Assesses evidence directness using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence directness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence directness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Overall Evidence Certainty
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Biomarker Sensitivity to Change
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Construct Validity
The degree to which construct supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000035
BEMO:2000035
42
5bd558ed006f260bde44545e9ab38951fc6c2e450b4d3cb07f929d1fcb479f4e
10
Assesses construct validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in construct validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating construct validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Quantification Precision
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Real-World Evidence Alignment
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Incremental Diagnostic Value
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of incremental diagnostic value.
bemo
BEMO:2000120
BEMO:2000120
127
8c0d8dc6983b3100a525f45a2c3c327af7d21cf2d155dc05a19b99e0d7f43578
8
Assesses incremental diagnostic value using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in incremental diagnostic value can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating incremental diagnostic value as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 473
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Allocation Concealment; Blinding Integrity; Performance Bias Risk", "Common Misinterpretations": "Treating baseline comparability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Baseline Comparability", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of baseline comparability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses baseline comparability using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in baseline comparability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
473
93fad13257c869638e96dc16ab4687d4b955a0d40d8d10736f49f12cebec65eb
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 57
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Diagnostic Biomarker Validity; Monitoring Biomarker Validity; Safety Biomarker Validity", "Common Misinterpretations": "Treating pharmacodynamic biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Pharmacodynamic Biomarker Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which pharmacodynamic biomarker supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses pharmacodynamic biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in pharmacodynamic biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
57
4eb497c772907b7927e32ee929fba9a804b1c75888594041caa61dfab977440a
E-value Strength
The magnitude and credibility of independent evidence supporting e-value.
bemo
BEMO:2000093
BEMO:2000093
100
8587ae49790e995db7d3bf84506f3bebdce4263532f6e5dcbd136229f5da442f
8
Assesses e-value strength using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in e-value strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating e-value strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Specialized / infrequent
Developing
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 125
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Verification Bias Risk; Incorporation Bias Risk; Patient Flow Integrity", "Common Misinterpretations": "Treating differential verification bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Differential Verification Bias Risk", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "The probability or degree that differential verification bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses differential verification bias risk using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in differential verification bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
125
2805e0d0838d99254f4675a5517e72c35cb3e1cb875fbc214edb03376568c858
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 188
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Positive-Control Performance; Vehicle-Control Validity; Orthogonal Validation", "Common Misinterpretations": "Treating negative-control performance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Negative-Control Performance", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of negative-control performance.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses negative-control performance using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in negative-control performance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
188
0095c2122f8f2e2a2aa6070b8c33f6576321db691574400350300c93ce956830
Computation profile for Pharmacodynamic Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Recovery
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of recovery.
bemo
BEMO:2000295
BEMO:2000295
302
5e396bd1803ce209b22910e5705b5bf0dc26bd235bba5ac5c904b2514e3f7ea8
8
Assesses recovery using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in recovery can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating recovery as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Selective Nonreporting Risk
The probability or degree that selective nonreporting introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000163
BEMO:2000163
170
8c42609174eeba28064b2f75bd158a40f97b17f9cdaa32aed81219eddf87ef89
8
Assesses selective nonreporting risk using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in selective nonreporting risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating selective nonreporting risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Biological Plausibility and Mechanism metric
Category of biomedical evidence metrics concerned with biological plausibility and mechanism.
bemo
BEMO:1100001
Candidate
Computation profile for Chain-of-Custody Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Generalizability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Processing Delay Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of processing delay control.
bemo
BEMO:2000076
BEMO:2000076
83
575df935e8ca0edd1975c05772eaa99f38b63f196b0ee598cef40ebcb4554aaf
8
Assesses processing delay control using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in processing delay control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating processing delay control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Orthogonal Validation
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of orthogonal validation.
bemo
BEMO:2000182
BEMO:2000182
189
86d52a57d493edc61e94352ed7d2d8f4feaaa492b37b0afb047315bb8706c00e
8
Assesses orthogonal validation using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in orthogonal validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating orthogonal validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 84
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "DNA Integrity; Microbial Contamination; Chain-of-Custody Integrity", "Common Misinterpretations": "Treating protein integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Protein Integrity", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of protein integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses protein integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in protein integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
84
f776649240f851879898d488bb90516491936e0936f6f03ac5946be6c8b62a20
limitations
Preserves the source limitations.
BEMO:3200007
Proteomics and Metabolomics metric
Category of biomedical evidence metrics concerned with proteomics and metabolomics.
bemo
BEMO:1100014
Candidate
Randomization in Experimental Allocation
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of randomization in experimental allocation.
bemo
BEMO:2000184
BEMO:2000184
191
5cc94b6085b9c596d4f74c18712e9f8177db4335ab703b2218b9c3a63209c113
10
Assesses randomization in experimental allocation using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in randomization in experimental allocation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating randomization in experimental allocation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 348
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "No-Observed-Adverse-Effect Level Robustness; Benchmark Dose Reliability; Toxicological Mode-of-Action Support", "Common Misinterpretations": "Treating lowest-observed-adverse-effect level robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Lowest-Observed-Adverse-Effect Level Robustness", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of lowest-observed-adverse-effect level robustness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses lowest-observed-adverse-effect level robustness using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in lowest-observed-adverse-effect level robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
348
cd00c7765086fbbe65040e649de98f22bdf10298c91d8da0f6e7300a18f0ba9f
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 339
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Genotoxicity Evidence Strength; Reproductive Toxicity Evidence Strength; Developmental Toxicity Evidence Strength", "Common Misinterpretations": "Treating carcinogenicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Carcinogenicity Evidence Strength", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The magnitude and credibility of independent evidence supporting carcinogenicity evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses carcinogenicity evidence strength using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in carcinogenicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
339
a171459b41fd9ad0436588a80669a48227118c77c3d87194654b56ab6dfa6ad7
Computation profile for Experimental Unit Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Cutoff Validity
The degree to which cutoff supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000278
BEMO:2000278
285
f74e484a95fd67b9bd73e0f0c112bc85ddaf01e7476f82b9c24c9382fe2bfe48
10
Assesses cutoff validity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in cutoff validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cutoff validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Selective Outcome Reporting Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 376
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Derivatization Efficiency; Pathway Enrichment Robustness; Cross-Omics Concordance", "Common Misinterpretations": "Treating metabolic feature reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Metabolic Feature Reproducibility", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The degree to which metabolic feature reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses metabolic feature reproducibility using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in metabolic feature reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
376
521bf8b7c1bbe892685e95c05cc7d63a05c4e26119e9951d705ff460b3bb99d1
Computation profile for Negative-Control Validation
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Negative-Result Reporting
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of negative-result reporting.
bemo
BEMO:2000419
BEMO:2000419
426
f6803e55dc7d7550d62e5bed1f8773c7da7d0266b37145da5354a96ceaf769c2
8
Assesses negative-result reporting using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in negative-result reporting can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating negative-result reporting as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Diagnostic Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Limit of Detection
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of limit of detection.
bemo
BEMO:2000285
BEMO:2000285
292
32751f0c3f93a66c901d3c5cea950ee45482dac2a524a6e4ed0ddc20000649d1
8
Assesses limit of detection using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in limit of detection can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating limit of detection as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 134
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Negative Predictive Value; Negative Likelihood Ratio; Diagnostic Odds Ratio", "Common Misinterpretations": "Treating positive likelihood ratio as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Positive Likelihood Ratio", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of positive likelihood ratio.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.", "Units or Scale (if applicable)": "Ratio scale; null typically 1", "What It Measures": "Assesses positive likelihood ratio using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in positive likelihood ratio can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
134
65ab6bc3c7bfa89dbab942e4977888e1ec5d3ea554c360ceffa76f4daab1fd7e
Computation profile for Environmental Standardization
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 387
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Protein Inference Reliability; Sequence Coverage; Quantification Precision", "Common Misinterpretations": "Treating proteome coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Proteome Coverage", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The proportion and representativeness of the relevant proteome captured by the evidence or measurement process.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses proteome coverage using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in proteome coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
387
5fec1390446b523e8b6072b1babcb11902375d326650ad448b800e258e0a8401
Computation profile for Ion Suppression Assessment
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Conceptual Replication Success
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of conceptual replication success.
bemo
BEMO:2000389
BEMO:2000389
396
45fcdc0cd9da1d73796f886d24683b4f40ead7f7a14a06e84d39631b66a910ad
10
Assesses conceptual replication success using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in conceptual replication success can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating conceptual replication success as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Confounding Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Dose Proportionality
0.1.0
Apply a validated analyte- and method-specific measurement procedure with calibration and quality control.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","unitRef":"REQUIRED"}
specimen_or_material; measurement_procedure; calibration_reference; quality_control_results; unit
replicate_measurements; environmental_conditions; instrument_version
xsd:decimal
Method- and analyte-specific physical units
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required.
Source maturity: Established; BEMO computation profile requires independent validation.
Vehicle-Control Validity
The degree to which vehicle-control supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000190
BEMO:2000190
197
b1aee501d417f186bec79f011141586c66408c5b9c272dd6db3a4fa3a34f419d
10
Assesses vehicle-control validity using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in vehicle-control validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating vehicle-control validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 384
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Peptide Identification Confidence; False Discovery Rate Control; Peptide-Spectrum Match Quality", "Common Misinterpretations": "Treating protein identification confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Protein Identification Confidence", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The justified degree of certainty assigned to protein identification given the quantity, quality, consistency, and limitations of supporting evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses protein identification confidence using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in protein identification confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
384
86cf8ad175ebd274f625e908935f946750c2c41a33f40e7a8dc33224d8db7d28
Computation profile for Protein Inference Reliability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Microbial Contamination
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of microbial contamination.
bemo
BEMO:2000072
BEMO:2000072
79
a7e7bf9a47cefd808a5ba122413c735b8b095ffc84592451f9857f4fe080e389
8
Assesses microbial contamination using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in microbial contamination can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating microbial contamination as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Evidence Freshness
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Single-Cell Viability
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Evidence Robustness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 330
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Pathway Enrichment Consistency; Knowledge-Graph Evidence Completeness; Knowledge-Graph Provenance Quality", "Common Misinterpretations": "Treating pathway topology support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Pathway Topology Support", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The magnitude and credibility of independent evidence supporting pathway topology.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses pathway topology support using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in pathway topology support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
330
f408d8b817c9e3179698dad395d8e94ea187b92ae192e0ab52e026106d82f40a
Computation profile for Context Sensitivity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Mature; BEMO computation profile requires independent validation.
Validated Measurement Procedure
Controlled BEMO ComputationMode value: ValidatedMeasurementProcedure.
BEMO:4000012
ValidatedMeasurementProcedure
Computation profile for Benchmark Dose Reliability
0.1.0
Apply a validated analyte- and method-specific measurement procedure with calibration and quality control.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","unitRef":"REQUIRED"}
specimen_or_material; measurement_procedure; calibration_reference; quality_control_results; unit
replicate_measurements; environmental_conditions; instrument_version
xsd:decimal
Method- and analyte-specific physical units
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 388
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Quantification Precision; Dynamic Range Coverage; Missing-Value Burden", "Common Misinterpretations": "Treating quantification accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Quantification Accuracy", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The closeness of quantification to the accepted reference or true value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses quantification accuracy using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in quantification accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
388
5205abd1da59c18e5b7272ce8586c2ec2bee590ad36aa3289ce71f4aa4c17c0b
Single-Cell Viability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of single-cell viability.
bemo
BEMO:2000262
BEMO:2000262
269
faf5bedb35b3c8e6f9b571b29c9788bb087192d600f7d5646da266201f76f5dd
8
Assesses single-cell viability using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in single-cell viability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating single-cell viability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 389
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Sequence Coverage; Quantification Accuracy; Dynamic Range Coverage", "Common Misinterpretations": "Treating quantification precision as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Quantification Precision", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The closeness of repeated estimates or measurements and the narrowness of uncertainty around quantification.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses quantification precision using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in quantification precision can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
389
6e5af7c1fffa7603da0abd7518a670fd09cade5fe3e47c8417886f91e6b4e07a
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 409
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Data Availability; Researcher-Degrees-of-Freedom Sensitivity; Multiverse Analysis Robustness", "Common Misinterpretations": "Treating random-seed stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Random-Seed Stability", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of random-seed stability.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses random-seed stability using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in random-seed stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
409
c1a5a2e956afa3ff1606fb1893282ee7b8d5d3733fbd04798630fe0746f86074
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 292
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Analytical Specificity; Limit of Quantification; Linearity", "Common Misinterpretations": "Treating limit of detection as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Limit of Detection", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of limit of detection.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Method- and analyte-specific physical units", "What It Measures": "Assesses limit of detection using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in limit of detection can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
292
32751f0c3f93a66c901d3c5cea950ee45482dac2a524a6e4ed0ddc20000649d1
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 254
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Variant Call Quality; Allelic Balance; Strand Bias", "Common Misinterpretations": "Treating genotype quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Genotype Quality", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of genotype quality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses genotype quality using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in genotype quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
254
6b3ff41c46f655393bf02a8730b778b41b8fad3c6f62302d7543aef197484ea3
Computation profile for Exposure–Response Relationship
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
computation protocol version
Version of the executed protocol.
BEMO:3100028
Computation profile for Endpoint Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Cellularity Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Cross-Omics Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Developing; BEMO computation profile requires independent validation.
Rubric Or Model Required
Controlled BEMO ComputationReadinessStatus value: RubricOrModelRequired.
BEMO:4000015
RubricOrModelRequired
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 237
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Variant Pathogenicity Evidence Strength; Segregation Evidence Strength; De Novo Evidence Strength", "Common Misinterpretations": "Treating population frequency compatibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Population Frequency Compatibility", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of population frequency compatibility.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses population frequency compatibility using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in population frequency compatibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
237
da2b6bc79f60e4e8121c9a5cd9a562fbe0d255feef3ecb78f1de13bffe50a0f4
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 312
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Cross-Omics Integration Coherence; Cross-Layer Directional Concordance; Latent-Factor Stability", "Common Misinterpretations": "Treating cross-omics replication as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Developing", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Cross-Omics Replication", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cross-omics replication.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses cross-omics replication using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cross-omics replication can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
312
5b9be25d7af680eef194d1d2e393dbdeaaf607f62c516cb41c75f5dfd1ba50b9
metric validation record
An information content entity documenting validation design, benchmark, results, scope, and status for a metric or computation profile.
bemo
BEMO:0000104
Candidate
Computation profile for Surrogate Endpoint Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Preanalytical Robustness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 119
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Calibration Slope; Observed-to-Expected Ratio; Prognostic Added Value", "Common Misinterpretations": "Treating calibration-in-the-large as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Calibration-in-the-Large", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of calibration-in-the-large.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses calibration-in-the-large using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in calibration-in-the-large can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
119
4a117ceb82993be69cf387d42f6626e25d4ff682e98eb458d01d55a60ca6f29d
Diagnostic Biomarker Validity
The degree to which diagnostic biomarker supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000040
BEMO:2000040
47
b87f6768315d9d24a6505cd78867a8b0dc9310fbd10aa04b46856994c12562e5
10
Assesses diagnostic biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in diagnostic biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating diagnostic biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Peptide Identification Confidence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 406
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Intralaboratory Repeatability; Result Reproducibility; Inferential Reproducibility", "Common Misinterpretations": "Treating method reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Method Reproducibility", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The degree to which method reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses method reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in method reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
406
b36c4ba0356ae6813b5bb255ebc6934635b0cbf83c44326cfea7019592803c5e
Variance Estimation Validity
The degree to which variance estimation supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000462
BEMO:2000462
469
43a84aea1ea6cc79814a3121476a165954a222aece0c751b1ff5414b40efbe28
10
Assesses variance estimation validity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in variance estimation validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating variance estimation validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 341
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Bioavailability; Time–Concentration Profile Adequacy; Metabolite Coverage", "Common Misinterpretations": "Treating dose proportionality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Dose Proportionality", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of dose proportionality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Method- and analyte-specific physical units", "What It Measures": "Assesses dose proportionality using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in dose proportionality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
341
6d7bafb114f62ee059333139662f0dfa2a0fb0aaab829ac8fcdb03ed2fbd0708
Computation profile for Loss-of-Function Mechanism Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Limit of Quantification
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of limit of quantification.
bemo
BEMO:2000286
BEMO:2000286
293
e110a21f976b308d8a01ff42efac55157e3af426e1182c74cf700e2fb4ce3595
8
Assesses limit of quantification using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in limit of quantification can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating limit of quantification as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Biomarker Sensitivity to Change
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biomarker sensitivity to change.
bemo
BEMO:2000032
BEMO:2000032
39
80beed232e03f65da1843b394120d8baef7fb4428401374caac926260837e2b2
9
Assesses biomarker sensitivity to change using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in biomarker sensitivity to change can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biomarker sensitivity to change as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 465
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Model Fit; Distributional Assumption Adequacy; Variance Estimation Validity", "Common Misinterpretations": "Treating residual diagnostics adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Residual Diagnostics Adequacy", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The extent to which residual diagnostics is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses residual diagnostics adequacy using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in residual diagnostics adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
465
1193234e513641f836d7ad506af2e69fb2fcba17a604fff5710f7b53f487ad89
has validation record
Connects a metric or computation profile to a validation record.
BEMO:3000014
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 162
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Consensus Strength; Evidence Completeness; Evidence Coverage", "Common Misinterpretations": "Treating evidence sufficiency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Sufficiency", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence sufficiency.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses evidence sufficiency using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence sufficiency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
162
516f7f59fc37d19455d99f12f79cf9f0baf5cd671807c3364f013a016d7ec68a
Incorporation Bias Risk
The probability or degree that incorporation bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000119
BEMO:2000119
126
00ceeb2a9486ba76c68ad0f07ad3fe11fadf573e3b3b3fb37898281142b34a91
10
Assesses incorporation bias risk using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in incorporation bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating incorporation bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
metric governance record
An information content entity documenting ownership, review, approval, lifecycle, version, and change history.
bemo
BEMO:0000105
Candidate
Model–Experiment Concordance
The degree of agreement in model–experiment across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000313
BEMO:2000313
320
433f97851837c60d123f8a332258d7d5b3872aa7c5564c19691b4c4c00871a8b
8
Assesses model–experiment concordance using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in model–experiment concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating model–experiment concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 413
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Multiverse Analysis Robustness", "Common Misinterpretations": "Treating specification-curve robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Specialized / infrequent", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Developing", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Specification-Curve Robustness", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of specification-curve robustness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses specification-curve robustness using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in specification-curve robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
413
290b2551a20c17c23c05ae84270ce7ac2c9d4a43b9817050cdbe8a353ce743ac
Computation profile for Causal Contrast Clarity
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Evidence Confidence
The justified degree of certainty assigned to evidence given the quantity, quality, consistency, and limitations of supporting evidence.
bemo
BEMO:2000144
BEMO:2000144
151
a6abc5b380d07d44d4ec1e48b63cf0e3bf7299ab7264d6edcc898627f3bd0952
10
Assesses evidence confidence using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 415
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Intervention Description Completeness; Eligibility Criteria Completeness; Recruitment Reporting Completeness", "Common Misinterpretations": "Treating comparator description completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Comparator Description Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of comparator description are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses comparator description completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in comparator description completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
415
a4617b71362b91773383e09e14a13eab2a1f342da234aa0f9bf7cc4400a6ff10
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 460
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Influential Observation Sensitivity; Interaction Assessment Adequacy; Overadjustment Bias Risk", "Common Misinterpretations": "Treating nonlinearity assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Nonlinearity Assessment", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of nonlinearity assessment.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses nonlinearity assessment using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in nonlinearity assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
460
aec401720de6539460ff9d49f62284a772b28790802eda1b3dc7c5fc2d67af3f
Computation profile for Single-Cell Ambient RNA Burden
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 291
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Repeatability; Reproducibility of Measurement; Analytical Sensitivity", "Common Misinterpretations": "Treating intermediate precision as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Intermediate Precision", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The closeness of repeated estimates or measurements and the narrowness of uncertainty around intermediate.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses intermediate precision using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in intermediate precision can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
291
543dc6e95f52dce066f892ad8d141154500ea9d0cf35d11564c84cc551f70c3b
Computation profile for Prognostic Calibration
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Prognostic Transportability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prognostic transportability.
bemo
BEMO:2000132
BEMO:2000132
139
c2720768cdf2be8ed6a3b935c205e9cb522ea1c958d036e680eee5e7cba4038e
8
Assesses prognostic transportability using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in prognostic transportability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating prognostic transportability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Monitoring Biomarker Validity
The degree to which monitoring biomarker supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000048
BEMO:2000048
55
cc605972d7119b4dabf5e0bc927d29fa05979e706703654bb3ea93c2ebe95461
10
Assesses monitoring biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in monitoring biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating monitoring biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 418
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Null-Result Interpretability; Reproducibility Information Completeness", "Common Misinterpretations": "Treating deviations-from-protocol transparency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Deviations-from-Protocol Transparency", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of deviations-from-protocol transparency.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses deviations-from-protocol transparency using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in deviations-from-protocol transparency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
418
6d08df9b94938aa4356ba3bbe963ef6fed2e0e499e09e30936b635d55a24d3c8
Steady-State Validity
The degree to which steady-state supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000327
BEMO:2000327
334
3a349724361b2712b35c94baaaa135ce8e313533d3edafc463242f1cddef411b
10
Assesses steady-state validity using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in steady-state validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating steady-state validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Temporal Biological Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Pathway Enrichment Robustness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Residual Diagnostics Adequacy
The extent to which residual diagnostics is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000458
BEMO:2000458
465
1193234e513641f836d7ad506af2e69fb2fcba17a604fff5710f7b53f487ad89
8
Assesses residual diagnostics adequacy using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in residual diagnostics adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating residual diagnostics adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Computation profile for Segregation Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Retention-Time Stability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 212
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Population Representativeness; Transportability; Generalizability", "Common Misinterpretations": "Treating sampling frame adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Sampling Frame Adequacy", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "The extent to which sampling frame is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses sampling frame adequacy using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in sampling frame adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
212
39b818e971de479787fde11a4523a1b0d31f5dcbc88f645b968cff60fa6bf381
Variant Classification Stability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of variant classification stability.
bemo
BEMO:2000234
BEMO:2000234
241
103d1dc365ad36d3caee0208333084352e814015d8cdbebbb9964106d3315f21
9
Assesses variant classification stability using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in variant classification stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating variant classification stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 422
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Outcome Definition Completeness; Comparator Description Completeness; Eligibility Criteria Completeness", "Common Misinterpretations": "Treating intervention description completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Intervention Description Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of intervention description are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses intervention description completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in intervention description completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
422
52ff9263e3651129913da102dae69c16126528e71dde3bf30595ca84301723b5
Peptide-Spectrum Match Quality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of peptide-spectrum match quality.
bemo
BEMO:2000375
BEMO:2000375
382
69f7898b8265a7daaa45bd90d4d5387c92da6521721d3c4901a5c208a63f0169
8
Assesses peptide-spectrum match quality using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in peptide-spectrum match quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating peptide-spectrum match quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Sequence Coverage
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Funding-Source Transparency
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of funding-source transparency.
bemo
BEMO:2000413
BEMO:2000413
420
35377c10ba740cabcf147ce4b6ab00e3a9c4087ca2ca810c337f7fe08223e6f9
8
Assesses funding-source transparency using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.
Material weakness in funding-source transparency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating funding-source transparency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Study report / dataset / evidence package
All biomedical study reports and data releases
EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS
https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Monitoring Biomarker Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Network Edge Confidence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 381
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Protein Identification Confidence; False Discovery Rate Control", "Common Misinterpretations": "Treating peptide identification confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Peptide Identification Confidence", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The justified degree of certainty assigned to peptide identification given the quantity, quality, consistency, and limitations of supporting evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses peptide identification confidence using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in peptide identification confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
381
7f78eede8aa05781dde3f12d7ea8111e1c0b552f5b722c402b0de5323e65b0ea
Genotype–Phenotype Concordance
The degree of agreement in genotype–phenotype across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000222
BEMO:2000222
229
baeee8e900a9ffef8f3ddd054b3a12f2233d89fc626f852fd31efc4b84b81685
8
Assesses genotype–phenotype concordance using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in genotype–phenotype concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.
Treating genotype–phenotype concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Molecular-Phenotypic Concordance
The degree of agreement in molecular-phenotypic across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000015
BEMO:2000015
22
81a78e88eaa62f6e2bb0ae947aba28fe819150032abcf9264f9205803cc7dfd1
8
Assesses molecular-phenotypic concordance using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in molecular-phenotypic concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating molecular-phenotypic concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Animal Model Predictive Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 271
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Allelic Balance; Reference Bias; Call-Rate Completeness", "Common Misinterpretations": "Treating strand bias as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Strand Bias", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of strand bias.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses strand bias using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in strand bias can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
271
6ac08afc8cf99d9cd5ee94bc247ff32edf260fc542fd5e27409d9c3b702afd66
uses metric
Connects a metric assessment to the metric used.
BEMO:3000002
Computation profile for Mechanistic Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Sample Identity Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Risk Assessment Computation
Controlled BEMO ComputationMode value: RiskAssessmentComputation.
BEMO:4000010
RiskAssessmentComputation
body of evidence
A defined collection of evidence records assembled to evaluate one or more biomedical claims.
bemo
BEMO:0000203
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 71
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Biospecimen Provenance Completeness; Warm Ischemia Control; Cold Ischemia Control", "Common Misinterpretations": "Treating collection procedure consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Collection Procedure Consistency", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "The degree of agreement in collection procedure across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses collection procedure consistency using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in collection procedure consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
71
c748a49d52923f74e8e31a0f69a4e91ee50970289d8b43c903a638e0573cc828
source row
One-based row number in the source worksheet.
BEMO:3100022
Code Availability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of code availability.
bemo
BEMO:2000387
BEMO:2000387
394
196f2dc58cfb1ea70e003d9904020fcfd1f4feec7308ace2e85d573eb71d3b5e
8
Assesses code availability using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in code availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating code availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Nonlinearity Assessment
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Dynamical Stability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 170
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Publication Bias Risk; Small-Study Effects; Study Heterogeneity", "Common Misinterpretations": "Treating selective nonreporting risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Selective Nonreporting Risk", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The probability or degree that selective nonreporting introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses selective nonreporting risk using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in selective nonreporting risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
170
8c42609174eeba28064b2f75bd158a40f97b17f9cdaa32aed81219eddf87ef89
Transport Condition Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of transport condition integrity.
bemo
BEMO:2000081
BEMO:2000081
88
81f1c97831f3ffbcc8b4e89c16604f987d3fa2baac750496d710981c68a60598
8
Assesses transport condition integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in transport condition integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating transport condition integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Single-Cell Doublet Burden
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 183
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Biological Replicate Adequacy; Technical Replicate Adequacy", "Common Misinterpretations": "Treating experimental unit validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Experimental Unit Validity", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The degree to which experimental unit supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses experimental unit validity using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in experimental unit validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
183
3d08eaaab51eb4d1095b9fb5e884110ba36de86e61615acc1661c581f985a6b1
Freeze–Thaw Burden
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of freeze–thaw burden.
bemo
BEMO:2000067
BEMO:2000067
74
40c5c886eacc67a9addeaa92c47d69bf9fbd4b1d2a0042a2efd3611fd5cebb4b
8
Assesses freeze–thaw burden using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in freeze–thaw burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating freeze–thaw burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Quantification Accuracy
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Causal Identifiability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Ontology Evidence-Code Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Transportability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Context-of-Use Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Mechanistic Coverage
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Population Frequency Compatibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Metabolite Coverage
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 318
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Knowledge-Graph Evidence Completeness; Entity Resolution Accuracy; Relation Evidence Strength", "Common Misinterpretations": "Treating knowledge-graph provenance quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Developing", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Knowledge-Graph Provenance Quality", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of knowledge-graph provenance quality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses knowledge-graph provenance quality using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in knowledge-graph provenance quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
318
50aaf6749314a5ad4694891f48d60c74feaa99da022b5cc91c17f4221f0696a2
Computation profile for Transcript Quantification Reliability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Negative Predictive Value
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of negative predictive value.
bemo
BEMO:2000123
BEMO:2000123
130
9f42e5362ad2e55445a028ec6d3cc9c062886d8c9272f983b0c1acec345502f7
8
Assesses negative predictive value using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in negative predictive value can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.
Treating negative predictive value as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 23
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Homeostatic Compensation Assessment; Systems-Level Emergence Support", "Common Misinterpretations": "Treating network context support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Network Context Support", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The magnitude and credibility of independent evidence supporting network context.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses network context support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in network context support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
23
d14cae7a1ae4e812a18da95395619423ade3837794dc32beea80cb11af0003f0
Computation profile for Repeatability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
missing data policy
Required missing-data handling policy.
BEMO:3100015
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 448
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Effect Magnitude; Confidence Interval Compatibility; Statistical Power", "Common Misinterpretations": "Treating estimate precision as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Estimate Precision", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The closeness of repeated estimates or measurements and the narrowness of uncertainty around estimate.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses estimate precision using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in estimate precision can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
448
9458d6e4371d11a80963339d3a9a6c5dcc187a36080cf18d964ecb5827c41820
has governance record
Connects an ontology entity to its governance record.
BEMO:3000015
Context Dependent Mixed Scale
Controlled BEMO ScaleType value: ContextDependentMixedScale.
BEMO:4000001
ContextDependentMixedScale
Computation profile for Analytical Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Response Biomarker Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 115
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Collider Bias Risk; Immortal-Time Bias Risk; Reverse-Causation Risk", "Common Misinterpretations": "Treating time-varying confounding control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Time-Varying Confounding Control", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of time-varying confounding control.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses time-varying confounding control using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in time-varying confounding control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
115
f46690f0cea7611ea5bb0e9a0e5c1b90d25ecad9c8f28f3083cc2a71b53ce7f0
Computation profile for Genotoxicity Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 476
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Period Effect Risk; Early Stopping Bias Risk; Adherence Integrity", "Common Misinterpretations": "Treating cluster recruitment bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Cluster Recruitment Bias Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that cluster recruitment bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses cluster recruitment bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cluster recruitment bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
476
a9882065e899d9f8a8f701c6291666b702a9ad5b769971c1d3cb596c9e39b960
Mapping Quality
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of mapping quality.
bemo
BEMO:2000250
BEMO:2000250
257
aef9eeffa971170a3db3642e806206bdf35806d1d8cc427e16ad9106f7e0207c
8
Assesses mapping quality using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in mapping quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.
Treating mapping quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Causal Inference metric
Category of biomedical evidence metrics concerned with causal inference.
bemo
BEMO:1100004
Candidate
Positive Predictive Value
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of positive predictive value.
bemo
BEMO:2000128
BEMO:2000128
135
89034784369de0142956f172eb3ca87ccad9b4cc69ac8a45b8a92adf9cc1731f
8
Assesses positive predictive value using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in positive predictive value can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.
Treating positive predictive value as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Residual Diagnostics Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Biomarker Qualification Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Exposure–Response Relationship
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of exposure–response relationship.
bemo
BEMO:2000337
BEMO:2000337
344
882d0346f9d5903a7e9b6a2b5ae992e9e24bd2be9c7e249168f2f9e157dc0354
8
Assesses exposure–response relationship using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in exposure–response relationship can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating exposure–response relationship as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Protocol Reproducibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 206
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Subgroup Consistency; Temporal Generalizability; Demographic Generalizability", "Common Misinterpretations": "Treating geographic consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Geographic Consistency", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "The degree of agreement in geographic across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses geographic consistency using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in geographic consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
206
22cb03e816656f6f9ad7a2acb4c32a01ffbc6b40d04a48011cef0a52134d504e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 253
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Sequencing Depth Adequacy; Mapping Quality; Read Quality", "Common Misinterpretations": "Treating genome coverage uniformity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Genome Coverage Uniformity", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "The proportion and representativeness of the relevant genome uniformity captured by the evidence or measurement process.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses genome coverage uniformity using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in genome coverage uniformity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
253
c86e8ba8cbfb79ea4dbc1a9a4cd5da76837e35c7f45f5966b51b9734760bdf1f
Computation profile for RNA Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
numeric metric value
Numeric metric assessment value.
BEMO:3100026
Computation profile for Correct Temporal Ordering
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Normalization Adequacy
The extent to which normalization is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000251
BEMO:2000251
258
518f55960ed771c22a0c382ff436c097ae34f4481b35b5a64e7f2cf8493778a4
8
Assesses normalization adequacy using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in normalization adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating normalization adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 269
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Single-Cell Ambient RNA Burden; Single-Cell Feature Detection Rate; Cell-Type Annotation Confidence", "Common Misinterpretations": "Treating single-cell viability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Single-Cell Viability", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of single-cell viability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses single-cell viability using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in single-cell viability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
269
faf5bedb35b3c8e6f9b571b29c9788bb087192d600f7d5646da266201f76f5dd
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 161
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Quality; Evidence Confidence; Evidence Stability", "Common Misinterpretations": "Treating evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Strength", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The magnitude and credibility of independent evidence supporting evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses evidence strength using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
161
d0050ec2d0692a0a792ec941fce4c076cf8613142effd9ef06a03cab99fa82fe
Overadjustment Bias Risk
The probability or degree that overadjustment bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000455
BEMO:2000455
462
20f394b3ef37b7c0b5d8426ebcce5fe9f9ef841bbb2622101174fbb411ab7957
10
Assesses overadjustment bias risk using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in overadjustment bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating overadjustment bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Computation profile for Matched-Sample Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Reproducibility, Replication, and Research Transparency metric
Metrics assessing reproducibility, independent replication, reporting completeness, and research transparency.
bemo
BEMO:1000004
Candidate
has computation profile
Connects a metric-bearing assessment or metric class restriction to its computation specification.
BEMO:3000001
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 21
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Biological Plausibility; Mechanistic Coverage; Mechanistic Completeness", "Common Misinterpretations": "Treating mechanistic support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Mechanistic Support", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The magnitude and credibility of independent evidence supporting mechanistic.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses mechanistic support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in mechanistic support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
21
d05eb179325de1077bb0eedc4d22a5f8c9cc2049d4ca5a5830b380984606cc90
Context Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of context sensitivity.
bemo
BEMO:2000193
BEMO:2000193
200
84e49bfba34c254503c96b181528bf2beab33e58bfa3d6d711b6afaa740ea566
9
Assesses context sensitivity using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in context sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating context sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
human-readable formula
Human-readable formula or protocol requirement.
BEMO:3100003
Computation profile for False Discovery Rate Control
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Dynamic Range
0.1.0
Apply a validated analyte- and method-specific measurement procedure with calibration and quality control.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","unitRef":"REQUIRED"}
specimen_or_material; measurement_procedure; calibration_reference; quality_control_results; unit
replicate_measurements; environmental_conditions; instrument_version
xsd:decimal
Method- and analyte-specific physical units
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 28
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Cross-Species Biological Concordance; Molecular-Phenotypic Concordance; Perturbational Validation", "Common Misinterpretations": "Treating phenotypic concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Phenotypic Concordance", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The degree of agreement in phenotypic across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses phenotypic concordance using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in phenotypic concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
28
d6b2622e1418a4a4c75dd4e74edee6d3fe7ea5ecac7eb6c453c614421e29173a
Area Under the Receiver Operating Characteristic Curve
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of area under the receiver operating characteristic curve.
bemo
BEMO:2000110
BEMO:2000110
117
5c4feaae4c2a565042108ac4cb807678f9be09f3e84e62f596afd3338e3ceec6
8
Assesses area under the receiver operating characteristic curve using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in area under the receiver operating characteristic curve can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating area under the receiver operating characteristic curve as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Evidence Precision
The closeness of repeated estimates or measurements and the narrowness of uncertainty around evidence.
bemo
BEMO:2000150
BEMO:2000150
157
fed12d9ff25f08fc6db399dc5bed2e8d719fc0db8bcefb33987a3ef3db40f55c
9
Assesses evidence precision using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence precision can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence precision as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Mature
Computation profile for Mediation Evidence Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Animal Model Face Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Reagent Lot Consistency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 11
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Temporal Biological Concordance; Phenotypic Concordance; Molecular-Phenotypic Concordance", "Common Misinterpretations": "Treating cross-species biological concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Cross-Species Biological Concordance", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The degree of agreement in cross-species biological across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses cross-species biological concordance using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cross-species biological concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
11
6f883b8ebe7aa3d7407688c8b2390b245a894c21f5c21ffef99fe634674c5fd3
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 433
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Prospective Registration; Prespecified Analysis Adherence", "Common Misinterpretations": "Treating protocol availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Protocol Availability", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of protocol availability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses protocol availability using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in protocol availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
433
6e6dfc0138818017108063ab02fea9da669c25ffc36ec988048eb058c538f84a
Endpoint Responsiveness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of endpoint responsiveness.
bemo
BEMO:2000043
BEMO:2000043
50
3cb1f45a15a5daa5f6b98c2d745a57440f42e5840e757c82a744e9f010f76df0
8
Assesses endpoint responsiveness using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in endpoint responsiveness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating endpoint responsiveness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Trial-Level Surrogacy
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 105
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Target Trial Emulation Fidelity; Negative-Control Validation; Dose–Response Support", "Common Misinterpretations": "Treating instrument validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Instrument Validity", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The degree to which instrument supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses instrument validity using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in instrument validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
105
e7033f91a59f115f9fba748b04cab66664ac151f4c8d653ea61aa823a8c6cce4
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 139
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Prognostic Added Value; Time-Dependent Discrimination; Competing-Risk Model Validity", "Common Misinterpretations": "Treating prognostic transportability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Prognostic Transportability", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prognostic transportability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses prognostic transportability using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in prognostic transportability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
139
c2720768cdf2be8ed6a3b935c205e9cb522ea1c958d036e680eee5e7cba4038e
Knowledge-Graph Evidence Completeness
The extent to which all scientifically necessary components of knowledge-graph evidence are present, documented, and evaluable.
bemo
BEMO:2000310
BEMO:2000310
317
082de1cac908d5d8b791ddee2974743d0e0eae71997e72080a5d978e931f8b6c
9
Assesses knowledge-graph evidence completeness using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in knowledge-graph evidence completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating knowledge-graph evidence completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Developing
Target Engagement Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of target engagement evidence.
bemo
BEMO:2000025
BEMO:2000025
32
a51c69d24f2bf9ef711f1d3b54d2086032245754bf1992c1b5999a88e3ca185f
8
Assesses target engagement evidence using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in target engagement evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating target engagement evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Entity Resolution Accuracy
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 214
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Outcome Applicability; Subgroup Consistency; Geographic Consistency", "Common Misinterpretations": "Treating spectrum representativeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Spectrum Representativeness", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of spectrum representativeness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses spectrum representativeness using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in spectrum representativeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
214
771c727a328004aa438f83d20410cd3600ce67ac5affa57fc6e9aaee55cb7ed9
Computation profile for Differential Follow-up Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Genome Coverage Uniformity
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Sequence Coverage
The proportion and representativeness of the relevant sequence captured by the evidence or measurement process.
bemo
BEMO:2000384
BEMO:2000384
391
f98e41911ad517a10150d0d24373bec15a553a9812e119c2e6a10081e9a89aaf
9
Assesses sequence coverage using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in sequence coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating sequence coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Instrument Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 53
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Biomarker Sensitivity to Change; Convergent Validity; Discriminant Validity", "Common Misinterpretations": "Treating known-groups validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Known-Groups Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which known-groups supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses known-groups validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in known-groups validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
53
8e195ef346deba393b7a3adaf4c77bfbd949dbf0f057cd9af26312568cb85421
Mechanistic Completeness
The extent to which all scientifically necessary components of mechanistic are present, documented, and evaluable.
bemo
BEMO:2000011
BEMO:2000011
18
70dc43fe10f07fb6cf3546a825ad424f5db2e21cc1b19a9f92b47d758bf660bd
9
Assesses mechanistic completeness using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in mechanistic completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating mechanistic completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
what it measures
Preserves the source explanation of what a metric measures.
BEMO:3200002
Ratio Scale
Controlled BEMO ScaleType value: RatioScale.
BEMO:4000005
RatioScale
Single-Cell Ambient RNA Burden
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of single-cell ambient rna burden.
bemo
BEMO:2000259
BEMO:2000259
266
1685a23504d9a42b00cb2ea9bdca0c2daeec5f3be11cc0d479a9acb5349c658a
8
Assesses single-cell ambient rna burden using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.
Material weakness in single-cell ambient rna burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating single-cell ambient rna burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies
MIAME; MINSEQE; STROBE-ME; GA4GH; HCA
https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Proteoform Identification Confidence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Positive Predictive Value
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Prognostic Added Value
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 187
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Environmental Standardization; Humane Endpoint Appropriateness; Attrition Accounting in Animal Studies", "Common Misinterpretations": "Treating intervention fidelity in animal studies as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Intervention Fidelity in Animal Studies", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of intervention fidelity in animal studies.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses intervention fidelity in animal studies using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in intervention fidelity in animal studies can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
187
f4ebc6af400b7084f27d64f43d25beed78acdab6caf4a03c3179cab231054034
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 425
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Statistical Methods Reporting Completeness; Funding-Source Transparency; Conflict-of-Interest Transparency", "Common Misinterpretations": "Treating missing-data reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Missing-Data Reporting Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of missing-data reporting are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses missing-data reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in missing-data reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
425
810e02862dfc4630d6f7a0d7bd830ccd305d847e2d0153cfdb863b68f1c2ece8
Computation profile for Clinical Relevance of Effect
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 202
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Demographic Generalizability; Care-Pathway Independence; Context Sensitivity", "Common Misinterpretations": "Treating disease-severity generalizability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Disease-Severity Generalizability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of disease-severity generalizability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses disease-severity generalizability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in disease-severity generalizability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
202
867ff07055e7596b98039def11c9ce9ca056af7a3a11134bb8f46a0bd5adc97a
Evidence Freshness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence freshness.
bemo
BEMO:2000149
BEMO:2000149
156
2dc879dd1ad8ed3ebf228b83983e7de95e06b9c72131b584ccaa785f551eea2b
8
Assesses evidence freshness using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence freshness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence freshness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 25
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Target Engagement Evidence; Off-Target Liability Evidence; Biological Gradient", "Common Misinterpretations": "Treating on-target specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "On-Target Specificity", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of on-target specificity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses on-target specificity using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in on-target specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
25
5b75004715c78ae57eb0dfbf1a6fe3b6212289ca6b06f39e35b36270d16b13e5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 18
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Mechanistic Coverage; Mechanistic Coherence; Mechanistic Specificity", "Common Misinterpretations": "Treating mechanistic completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Mechanistic Completeness", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The extent to which all scientifically necessary components of mechanistic are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses mechanistic completeness using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in mechanistic completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
18
70dc43fe10f07fb6cf3546a825ad424f5db2e21cc1b19a9f92b47d758bf660bd
Unmeasured Confounding Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of unmeasured confounding sensitivity.
bemo
BEMO:2000109
BEMO:2000109
116
50677da95ed69c14f1f1ec5786cb5c4d590d6e339c1a13fdedfd18d811ffc8f4
10
Assesses unmeasured confounding sensitivity using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in unmeasured confounding sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.
Treating unmeasured confounding sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 66
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Biospecimen Provenance Completeness; Collection Procedure Consistency", "Common Misinterpretations": "Treating biospecimen integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Biospecimen Integrity", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biospecimen integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses biospecimen integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biospecimen integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
66
b89a21b15f8ca110af0a2e2a28ae4ce7939819f7d62769c459c1cd5dc9b712cc
Analytical Sensitivity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of analytical sensitivity.
bemo
BEMO:2000270
BEMO:2000270
277
193a9fb566dfa18750368d0657f7d1e8c3de7251eea014f0d99b1b3e73df3afa
10
Assesses analytical sensitivity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in analytical sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating analytical sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
unit text
Human-readable unit or scale.
BEMO:3100009
Computation profile for Negative Predictive Value
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 444
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Calibration of Statistical Predictions; Decision-Curve Net Benefit; Clinical Relevance of Effect", "Common Misinterpretations": "Treating discrimination of statistical predictions as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Discrimination of Statistical Predictions", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of discrimination of statistical predictions.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses discrimination of statistical predictions using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in discrimination of statistical predictions can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
444
cf100cd63561a5d7a167b9d561a196b9309ca048456a8705766e76a54a524fce
Confounding Risk
The probability or degree that confounding introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000088
BEMO:2000088
95
cebd58c93ed86b363870e131c5d23217522d15af0abb7eac8956ec81e8515565
10
Assesses confounding risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in confounding risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.
Treating confounding risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pathway-Level Support
The magnitude and credibility of independent evidence supporting pathway-level.
bemo
BEMO:2000019
BEMO:2000019
26
a1a0e0a377d9f949a45a0ad0a040b0443e29f24bfe701f3b015420b52cd21ab2
8
Assesses pathway-level support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in pathway-level support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating pathway-level support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
maturity of metric
Source-provided maturity classification.
BEMO:3200020
Mechanistic Causality Strength
The magnitude and credibility of independent evidence supporting mechanistic causality.
bemo
BEMO:2000009
BEMO:2000009
16
114851a566f6667271761fd50b701ab810f623c4ee0ecdabe88d5bbacc5941db
10
Assesses mechanistic causality strength using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in mechanistic causality strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.
Treating mechanistic causality strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
metric assessment
A measurement datum or structured assessment produced by applying a biomedical evidence metric to a defined evidence object using a specified computation profile.
bemo
BEMO:0000200
Candidate
Computation profile for Mechanistic Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 29
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Gain-of-Function Validation; Epistasis Support; Target Engagement Evidence", "Common Misinterpretations": "Treating rescue experiment support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Rescue Experiment Support", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The magnitude and credibility of independent evidence supporting rescue experiment.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses rescue experiment support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in rescue experiment support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
29
9f3de30fb4252544b664477f26ad879c47cffd18e4a902652a07ca5b9fa59ec2
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 201
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Temporal Generalizability; Disease-Severity Generalizability; Care-Pathway Independence", "Common Misinterpretations": "Treating demographic generalizability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Demographic Generalizability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of demographic generalizability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses demographic generalizability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in demographic generalizability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
201
e22341c32d6f7aad19740ed4b08e7f5ffc168107f1e04a79ba4ec60f6b92c463
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 403
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Experimental Reproducibility; Intralaboratory Repeatability; Method Reproducibility", "Common Misinterpretations": "Treating interlaboratory reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Interlaboratory Reproducibility", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The degree to which interlaboratory reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses interlaboratory reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in interlaboratory reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
403
0977ae8d7278730723c8d024585a95d1125aa9fe6da897fe977a079b172ddb48
Temporal Generalizability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of temporal generalizability.
bemo
BEMO:2000209
BEMO:2000209
216
dc35bc0fc74498ddd3acc1a47faf52ec96fba459b80f9800025d1a3d1e674317
8
Assesses temporal generalizability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.
Material weakness in temporal generalizability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating temporal generalizability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Clinical, epidemiologic, diagnostic, translational, and population studies
GRADE; QUADAS-2; CONSORT; STROBE
https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 233
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Hotspot/Functional-Domain Evidence; Splicing Evidence Strength; RNA Evidence Strength", "Common Misinterpretations": "Treating null-variant quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Null-Variant Quality", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of null-variant quality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses null-variant quality using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in null-variant quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
233
de2f4749e64a7014070a28c9c8587efaf4dbb9b3583c2ab6e78cd6826c00de34
Evidence Synthesis and Certainty metric
Metrics assessing the completeness, certainty, consistency, precision, and synthesis of biomedical evidence.
bemo
BEMO:1000001
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 52
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Surrogate Endpoint Validity; Trial-Level Surrogacy; Endpoint Validity", "Common Misinterpretations": "Treating individual-level surrogacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Individual-Level Surrogacy", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of individual-level surrogacy.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses individual-level surrogacy using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in individual-level surrogacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
52
4d46697f970727d741077ee1e1f9eeeaed4beee8d6ef96a339ca927181b91ef5
Toxicokinetic Concordance
The degree of agreement in toxicokinetic across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000357
BEMO:2000357
364
6cda2f6cb167ccc21d05523345b2fef382432ceed6ebbcbe84fa383cddc4562e
8
Assesses toxicokinetic concordance using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in toxicokinetic concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating toxicokinetic concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Candidate
Controlled BEMO LifecycleStatus value: Candidate.
BEMO:4000023
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 241
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Genotype–Phenotype Concordance; Conflicting Interpretation Burden", "Common Misinterpretations": "Treating variant classification stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Variant Classification Stability", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of variant classification stability.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses variant classification stability using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in variant classification stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
241
103d1dc365ad36d3caee0208333084352e814015d8cdbebbb9964106d3315f21
frequency of use
Source-provided frequency-of-use classification.
BEMO:3200019
Lower Is Better
Controlled BEMO Directionality value: LowerIsBetter.
BEMO:4000022
LowerIsBetter
Computation profile for Evidence Confidence
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 197
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Negative-Control Performance; Orthogonal Validation; Technical Artifact Exclusion", "Common Misinterpretations": "Treating vehicle-control validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Vehicle-Control Validity", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The degree to which vehicle-control supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses vehicle-control validity using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in vehicle-control validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
197
b1aee501d417f186bec79f011141586c66408c5b9c272dd6db3a4fa3a34f419d
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 47
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Prognostic Biomarker Validity; Pharmacodynamic Biomarker Validity; Monitoring Biomarker Validity", "Common Misinterpretations": "Treating diagnostic biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Diagnostic Biomarker Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which diagnostic biomarker supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses diagnostic biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in diagnostic biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
47
b87f6768315d9d24a6505cd78867a8b0dc9310fbd10aa04b46856994c12562e5
Computation profile for Perturbation Prediction Accuracy
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Blinding Integrity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of blinding integrity.
bemo
BEMO:2000467
BEMO:2000467
474
4e84b980ded4e833e93013ba317d5857e2211ad9b612e4b77eb9308cf8244573
8
Assesses blinding integrity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.
Material weakness in blinding integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating blinding integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies
CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Population Frequency Compatibility
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of population frequency compatibility.
bemo
BEMO:2000230
BEMO:2000230
237
da2b6bc79f60e4e8121c9a5cd9a562fbe0d255feef3ecb78f1de13bffe50a0f4
8
Assesses population frequency compatibility using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in population frequency compatibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating population frequency compatibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Calibration Slope
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of calibration slope.
bemo
BEMO:2000111
BEMO:2000111
118
44bfcfec74d2b98de3507998b5e63f6506912c93210fea4df7a678aed4417a75
8
Assesses calibration slope using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in calibration slope can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating calibration slope as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Comparative Test Accuracy
The closeness of comparative test to the accepted reference or true value.
bemo
BEMO:2000113
BEMO:2000113
120
01f4b65430acfaca391e550bb1550450d424a1dd47aebe4ed0f6dd9d60fe28ab
8
Assesses comparative test accuracy using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in comparative test accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating comparative test accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Age Appropriateness
The extent to which age is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000166
BEMO:2000166
173
6f50922de3489bf44a56c621fcc8933973821a524b4c3ba13cb746276b4a9bea
8
Assesses age appropriateness using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in age appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating age appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 436
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Eligibility Criteria Completeness; Participant Flow Completeness; Harms Reporting Completeness", "Common Misinterpretations": "Treating recruitment reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Recruitment Reporting Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of recruitment reporting are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses recruitment reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in recruitment reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
436
d9e587b4cd42324cc1668fb835e848ce02fea72bac0af2f37bcc6e0a40cbc327
Computation profile for DNA Integrity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Posterior Probability Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Preanalytical Robustness
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of preanalytical robustness.
bemo
BEMO:2000293
BEMO:2000293
300
b76350bac6098346689193598ba9a67fbdcf0cc16c9a9abb5e49d8238a97798b
10
Assesses preanalytical robustness using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in preanalytical robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating preanalytical robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 15
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Perturbational Validation; Gain-of-Function Validation; Rescue Experiment Support", "Common Misinterpretations": "Treating loss-of-function validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Loss-of-Function Validation", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of loss-of-function validation.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses loss-of-function validation using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in loss-of-function validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
15
7b09e166eefa32f832c89264de2b3521f93b540fbc674d78fc625fd792779d09
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 324
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Latent-Factor Stability; Network Edge Confidence; Network Node Confidence", "Common Misinterpretations": "Treating network reconstruction robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Network Reconstruction Robustness", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of network reconstruction robustness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses network reconstruction robustness using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in network reconstruction robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
324
0b4f2bdfa3ccc3b3abdb67b1f3be15b6d8479743942152580ad3f1929db67d5f
Metabolite Identification Confidence
The justified degree of certainty assigned to metabolite identification given the quantity, quality, consistency, and limitations of supporting evidence.
bemo
BEMO:2000371
BEMO:2000371
378
dd08774f62d4c8036bc0f7c590a94eca58c7ca0613c38dec0e386cab59cc0dbd
10
Assesses metabolite identification confidence using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in metabolite identification confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.
Treating metabolite identification confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Calibration Slope
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 30
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Tissue Specificity; Temporal Biological Concordance; Cross-Species Biological Concordance", "Common Misinterpretations": "Treating spatial biological concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Spatial Biological Concordance", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The degree of agreement in spatial biological across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses spatial biological concordance using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in spatial biological concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
30
6c08ba26b1257f37b62212d7633a9e83166a0f74706ae3ffbb9de90cd98613f7
Deprecated
Controlled BEMO LifecycleStatus value: Deprecated.
BEMO:4000025
Deprecated
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 68
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Tumor Purity; Necrosis Burden; Hemolysis Burden", "Common Misinterpretations": "Treating cellularity adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Cellularity Adequacy", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "The extent to which cellularity is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses cellularity adequacy using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cellularity adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
68
de5efa73121b798bb0e1a53bddae9ef743d4a522e6df52007a5b7569bc63e020
Clinical Validity
The degree to which clinical supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000034
BEMO:2000034
41
628005d99d3e18a7b100ce400c64fc5218e46beeaa2cfec80dbac2445270cead
10
Assesses clinical validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in clinical validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating clinical validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Site-to-Site Assay Portability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 270
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Cell-Type Annotation Confidence; Cross-Platform Genomic Concordance", "Common Misinterpretations": "Treating spatial transcriptomic registration accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Spatial Transcriptomic Registration Accuracy", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "The closeness of spatial transcriptomic registration to the accepted reference or true value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses spatial transcriptomic registration accuracy using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in spatial transcriptomic registration accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
270
e61d8e9d42dcb8b644bc33652bdf4f396f77ace1c2d848fb684d0c722e5a1598
Susceptibility/Risk Biomarker Validity
The probability or degree that susceptibility/risk biomarker validity introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000056
BEMO:2000056
63
596927f557190db71a0b686158bde6134cafa9cdb9c1876784270d91a80452f5
10
Assesses susceptibility/risk biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in susceptibility/risk biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating susceptibility/risk biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 273
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Sex Concordance; Genotype Quality; Allelic Balance", "Common Misinterpretations": "Treating variant call quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Variant Call Quality", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of variant call quality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses variant call quality using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in variant call quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
273
45320c5b96fc37b8148ee2721453c92e6bec79491a4ed2d2f6fedbc2cb99cb01
False Discovery Rate Control
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of false discovery rate control.
bemo
BEMO:2000363
BEMO:2000363
370
4059adb92dbe8fea43e57c309bf3c42531450ae52d814c6f79932af37f2f97a2
8
Assesses false discovery rate control using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in false discovery rate control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating false discovery rate control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Pathway Topology Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 220
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Allelic Evidence Strength; Case-Control Evidence Strength; Functional Variant Evidence", "Common Misinterpretations": "Treating case-level evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Case-Level Evidence Strength", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The magnitude and credibility of independent evidence supporting case-level evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses case-level evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in case-level evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
220
e9e681d8a3c5b4ca82348a6cd8fe70f3f48bd8351d482c3aa492a6712873c3db
Reproducibility and Replication metric
Category of biomedical evidence metrics concerned with reproducibility and replication.
bemo
BEMO:1100015
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 264
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Genome Coverage Uniformity; Mapping Quality", "Common Misinterpretations": "Treating sequencing depth adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Sequencing Depth Adequacy", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "The extent to which sequencing depth is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses sequencing depth adequacy using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in sequencing depth adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
264
719d5f0d3157f75114503dd8e8b0e75aa481929ab009d6c25b68024c7b5f71a0
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 328
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Parameter Identifiability; Structural Identifiability; Practical Identifiability", "Common Misinterpretations": "Treating parameter sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Parameter Sensitivity", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of parameter sensitivity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses parameter sensitivity using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in parameter sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
328
9735c06335e80bafcbeea74ccb7b405e18e8684330b82448424f6e83ef3974ba
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 362
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Efficacy Reproducibility; Safety Margin Evidence; No-Observed-Adverse-Effect Level Robustness", "Common Misinterpretations": "Treating therapeutic window evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Therapeutic Window Evidence", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of therapeutic window evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses therapeutic window evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in therapeutic window evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
362
34f5d4730ddbb7e269fc6713d96c49c0fba2d2f738ddbfb47d339a6ed5f8cf6e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 45
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Known-Groups Validity; Discriminant Validity; Criterion Validity", "Common Misinterpretations": "Treating convergent validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Convergent Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which convergent supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses convergent validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in convergent validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
45
4ad4866845885f677cad74042f65b0303f2d26597ea694b76dfed4bc71fedd68
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 219
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Case-Level Evidence Strength; Functional Variant Evidence; Computational Variant Evidence", "Common Misinterpretations": "Treating case-control evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Case-Control Evidence Strength", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The magnitude and credibility of independent evidence supporting case-control evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses case-control evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in case-control evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
219
a2868c878aff5e68e6f9aeb87fd5a98fd776f1ca106d2e9dbad34a8cc81ecfd8
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 481
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Performance Bias Risk; Attrition Bias Risk; Selective Outcome Reporting Risk", "Common Misinterpretations": "Treating detection bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Detection Bias Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that detection bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses detection bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in detection bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
481
98c36b49fe475eec66359fae3b2440fc1429c378e2bbe5896e082ec203855535
Computation profile for Biological Replicate Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 252
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Read Quality; Library Complexity; Contamination Burden", "Common Misinterpretations": "Treating duplicate read burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Duplicate Read Burden", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of duplicate read burden.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses duplicate read burden using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in duplicate read burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
252
c1d6b63aea87c65750f96b5e3a771c6866843a2e5ba7e8081dd07af03bc0e9c5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 394
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Material Availability; Data Availability; Random-Seed Stability", "Common Misinterpretations": "Treating code availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Code Availability", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of code availability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses code availability using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in code availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
394
196f2dc58cfb1ea70e003d9904020fcfd1f4feec7308ace2e85d573eb71d3b5e
Dose–Response Support
The magnitude and credibility of independent evidence supporting dose–response.
bemo
BEMO:2000092
BEMO:2000092
99
75f662d9e2ee059c317733c8704b2f0f1f234146f1b2920ed572ec892b5a2d41
8
Assesses dose–response support using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.
Material weakness in dose–response support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating dose–response support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Randomized and observational etiologic studies, natural experiments, target-trial emulations
ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9
https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Read Quality
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 446
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Estimate Precision; Confidence Interval Compatibility", "Common Misinterpretations": "Treating effect magnitude as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Effect Magnitude", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of effect magnitude.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses effect magnitude using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in effect magnitude can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
446
e21896fd55cc8da1ecb273cc15de195832d87a4d4554d2b68b3a9ce060442d84
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 102
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Unmeasured Confounding Sensitivity; Positivity Adequacy; Consistency Assumption Plausibility", "Common Misinterpretations": "Treating exchangeability plausibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Exchangeability Plausibility", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of exchangeability plausibility.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses exchangeability plausibility using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in exchangeability plausibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
102
7f4004986e0cd5fd0b8462346aec167d3be6e60fe67bb8a5838de7a005fd58d1
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 391
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Proteome Coverage; Quantification Precision; Quantification Accuracy", "Common Misinterpretations": "Treating sequence coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Sequence Coverage", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The proportion and representativeness of the relevant sequence captured by the evidence or measurement process.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses sequence coverage using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in sequence coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
391
f98e41911ad517a10150d0d24373bec15a553a9812e119c2e6a10081e9a89aaf
Approved
Controlled BEMO ApprovalStatus value: Approved.
BEMO:4000029
Approved
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 91
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "E-value Strength; Falsification Test Support", "Common Misinterpretations": "Treating assumption sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Assumption Sensitivity", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of assumption sensitivity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses assumption sensitivity using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in assumption sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
91
b75d62759fbd8f126f27997543f0f46dde1cb8c2484d156a08c9925c9476aa63
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 340
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Reproductive Toxicity Evidence Strength; Immunotoxicity Evidence Strength; Toxicokinetic Concordance", "Common Misinterpretations": "Treating developmental toxicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Developmental Toxicity Evidence Strength", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The magnitude and credibility of independent evidence supporting developmental toxicity evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses developmental toxicity evidence strength using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in developmental toxicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
340
f5636ed02b8746eb48d98671b027c695a8940399bff60ebeae8ca955d99bae6a
Operator Variability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of operator variability.
bemo
BEMO:2000291
BEMO:2000291
298
ae9e50b8f8857543ecb1eb315c23fc88dd3a028cfd40ee4377e1039431544ca0
8
Assesses operator variability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in operator variability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating operator variability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 67
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Biospecimen Integrity; Collection Procedure Consistency; Warm Ischemia Control", "Common Misinterpretations": "Treating biospecimen provenance completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Biospecimen Provenance Completeness", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "The extent to which all scientifically necessary components of biospecimen provenance are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses biospecimen provenance completeness using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biospecimen provenance completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
67
c859051c1fe5f3cb60cda0bb4dcf804459be135ee18bb1bb44d89345dea50253
Computation profile for Ecological Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Module Stability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of module stability.
bemo
BEMO:2000314
BEMO:2000314
321
191acd0c1abb112365dc2e82b9a1c639642c0798eb52983d078ac92e3f0fbeec
9
Assesses module stability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in module stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating module stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Recovery
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Statistical Validity and Inference metric
Category of biomedical evidence metrics concerned with statistical validity and inference.
bemo
BEMO:1100017
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 137
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Prognostic Discrimination; Calibration Slope; Calibration-in-the-Large", "Common Misinterpretations": "Treating prognostic calibration as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Prognostic Calibration", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prognostic calibration.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses prognostic calibration using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in prognostic calibration can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
137
bddeb4ff61529649f6f5d81ae9a9ab5588a6f6aed10ca70d0d400119d6c2e940
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 168
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Between-Study Variance; Cumulative Evidence Stability; Information Size Adequacy", "Common Misinterpretations": "Treating prediction interval adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Prediction Interval Adequacy", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The extent to which prediction interval is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses prediction interval adequacy using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in prediction interval adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
168
0acc1464b5d065a6f38744191dff4d933036cf918aeac3f16274b9cdb9dd1059
Computation profile for Bioavailability
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 34
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Cell-Type Specificity; Spatial Biological Concordance; Temporal Biological Concordance", "Common Misinterpretations": "Treating tissue specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Tissue Specificity", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of tissue specificity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses tissue specificity using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in tissue specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
34
e3cc6b49f3bcbfabd500e6222a7c3cc053aca55c824a6afddd04a9b448378651
Rubric Required
Controlled BEMO ComputationReadinessStatus value: RubricRequired.
BEMO:4000016
RubricRequired
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 278
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Analytical Sensitivity; Limit of Detection; Limit of Quantification", "Common Misinterpretations": "Treating analytical specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Analytical Specificity", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of analytical specificity.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses analytical specificity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in analytical specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
278
f6458dae6787eecc6dbfe775e26236c89d02f80c2f92e3575c7f9ddfa0e84b0e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 46
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Discriminant Validity; Construct Validity; Content Validity", "Common Misinterpretations": "Treating criterion validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Criterion Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which criterion supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses criterion validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in criterion validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
46
04af58cc2d88f851cf92e0d1960d29a74b3d8a5cf3525dda9579101a813732b2
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 490
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Selective Outcome Reporting Risk; Contamination Risk; Co-intervention Bias Risk", "Common Misinterpretations": "Treating protocol deviation risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Protocol Deviation Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that protocol deviation introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses protocol deviation risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in protocol deviation risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
490
7632dd2b67a9f407b6c186b3a54ef8714ff7082ed9c885957d4f06c0008e5f26
Evidence Triangulation Strength
The magnitude and credibility of independent evidence supporting evidence triangulation.
bemo
BEMO:2000156
BEMO:2000156
163
b9d887ba211f3aea4699afd087a0aa12fb8ca924209ee34ccda2bfd968c02a4a
8
Assesses evidence triangulation strength using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence triangulation strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence triangulation strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 417
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Conflict-of-Interest Transparency; Code-Sharing Transparency; Materials-and-Reagents Reporting Completeness", "Common Misinterpretations": "Treating data-sharing transparency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Data-Sharing Transparency", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of data-sharing transparency.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses data-sharing transparency using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in data-sharing transparency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
417
24a8774755f1b71359ef7a35f4a7813a446670f2da0ccc38d22836e1a84f1453
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 327
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Model–Experiment Concordance; Parameter Sensitivity; Structural Identifiability", "Common Misinterpretations": "Treating parameter identifiability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Parameter Identifiability", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of parameter identifiability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses parameter identifiability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in parameter identifiability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
327
e1eaab072c302dca725ea1fb92e565c0aead2c5a6ddd49490dbfa43ad163dd03
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 129
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Positive Likelihood Ratio; Diagnostic Odds Ratio; Area Under the Receiver Operating Characteristic Curve", "Common Misinterpretations": "Treating negative likelihood ratio as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Negative Likelihood Ratio", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of negative likelihood ratio.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.", "Units or Scale (if applicable)": "Ratio scale; null typically 1", "What It Measures": "Assesses negative likelihood ratio using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in negative likelihood ratio can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
129
6dc29889e79017542636c8b841aff4c0f5947b19dec51307e00a9c51cdc593cc
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 171
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Selective Nonreporting Risk; Study Heterogeneity; Between-Study Variance", "Common Misinterpretations": "Treating small-study effects as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Small-Study Effects", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of small-study effects.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Forest plots; heterogeneity statistics; tau-squared; prediction intervals; funnel plots; regression or selection models; sensitivity analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses small-study effects using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in small-study effects can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
171
ec5ee159f748d5035209ae456b71092c6724bc73af9b720bf70afe3ffc41f242
Computation profile for Endpoint Reliability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 355
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Exposure–Response Relationship; Pharmacokinetic Adequacy", "Common Misinterpretations": "Treating pharmacological target validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Pharmacological Target Validity", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The degree to which pharmacological target supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses pharmacological target validity using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in pharmacological target validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
355
2cff4752043ea892788f37639d922b1a249d83c13dfedf581ad1a4ad12ab95b8
Computation profile for Demographic Generalizability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Measurement Uncertainty
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Interference Susceptibility
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of interference susceptibility.
bemo
BEMO:2000283
BEMO:2000283
290
af7341a32f2656e5638dc5bce7ac126974e88c78554d2916a696247ebbe11d2b
8
Assesses interference susceptibility using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in interference susceptibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating interference susceptibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Evidence Sufficiency
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence sufficiency.
bemo
BEMO:2000155
BEMO:2000155
162
516f7f59fc37d19455d99f12f79cf9f0baf5cd671807c3364f013a016d7ec68a
8
Assesses evidence sufficiency using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in evidence sufficiency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating evidence sufficiency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 454
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Variance Estimation Validity; Missing-Data Sensitivity; Imputation Validity", "Common Misinterpretations": "Treating missing-data mechanism plausibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Missing-Data Mechanism Plausibility", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of missing-data mechanism plausibility.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses missing-data mechanism plausibility using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in missing-data mechanism plausibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
454
399e0c1d410c7ca1c5cc767b5afb9291915d1ee3db3b4d7314b2092c6d4765e1
Cross-Omics Concordance
The degree of agreement in cross-omics across measurements, studies, methods, populations, or biological levels.
bemo
BEMO:2000359
BEMO:2000359
366
447ad3207c1d5333015f6a5189727d810a3a517e4d6a4351f8dc28acbe93a5f0
8
Assesses cross-omics concordance using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.
Material weakness in cross-omics concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating cross-omics concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies
MIAPE; HUPO PSI; Metabolomics Standards
https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Developing
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 113
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Reverse-Causation Risk; Quantitative Bias Analysis Robustness; E-value Strength", "Common Misinterpretations": "Treating selection-on-survival bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Selection-on-Survival Bias Risk", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The probability or degree that selection-on-survival bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses selection-on-survival bias risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in selection-on-survival bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
113
ba5e2bde29a386ccb9cae3afe1f1c9476a350f65371bc1aebd58a3579e6e702c
De Novo Evidence Strength
The magnitude and credibility of independent evidence supporting de novo evidence.
bemo
BEMO:2000217
BEMO:2000217
224
e1a17a41d450c91691b405a890d0fa9af7ecd8fed12385f8124d85b912c16860
8
Assesses de novo evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in de novo evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating de novo evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Lipemia Burden
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 469
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Distributional Assumption Adequacy; Missing-Data Mechanism Plausibility; Missing-Data Sensitivity", "Common Misinterpretations": "Treating variance estimation validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Variance Estimation Validity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The degree to which variance estimation supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses variance estimation validity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in variance estimation validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
469
43a84aea1ea6cc79814a3121476a165954a222aece0c751b1ff5414b40efbe28
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 153
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Stability; Evidence Directness; Evidence Precision", "Common Misinterpretations": "Treating evidence consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Consistency", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The degree of agreement in evidence across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses evidence consistency using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
153
f74a279906eee7e4453803727494ade321b94fc1d6196fb1882daae7f1abcf75
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 184
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Age Appropriateness; Housing and Husbandry Control; Environmental Standardization", "Common Misinterpretations": "Treating genetic background control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Genetic Background Control", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of genetic background control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses genetic background control using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in genetic background control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
184
3f6791503c2895971802b5fb5573311999fc32cde51839dc771c56e4855ceb83
Organ-Specific Toxicity Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of organ-specific toxicity evidence.
bemo
BEMO:2000345
BEMO:2000345
352
277a918d22b52db9e79eabf5451f64320f342f326be1cc2382f16c69ec72f0a9
8
Assesses organ-specific toxicity evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in organ-specific toxicity evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating organ-specific toxicity evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 279
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Accuracy; Trueness", "Common Misinterpretations": "Treating analytical validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Analytical Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The degree to which analytical supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses analytical validity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in analytical validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
279
e4b1deae92fd2c549054f6f19f234dbf0b73ab8ff30ff5741db1c88b5618f6c3
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 459
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Equivalence Margin Validity", "Common Misinterpretations": "Treating noninferiority margin validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Noninferiority Margin Validity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The degree to which noninferiority margin supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses noninferiority margin validity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in noninferiority margin validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
459
5c494db875d0f3ad6860a3bba83e6ac01d9f571b40aa89617a37a42c1c489cab
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 160
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Confidence; Evidence Consistency; Evidence Directness", "Common Misinterpretations": "Treating evidence stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Evidence Stability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence stability.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses evidence stability using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
160
1f9ca5bc5be56489e2caf9d2f3eaaf37b12ae4271716853fb373d2db6aa8d657
Computation profile for Period Effect Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Interference Susceptibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 9
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Mechanistic Support; Mechanistic Coverage", "Common Misinterpretations": "Treating biological plausibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Biological Plausibility", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biological plausibility.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses biological plausibility using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biological plausibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
9
6cbc5f525d668cebb0928b8428848a625103a7b6492d0a2a618adb116f200bd3
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 83
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Transport Condition Integrity; Anatomical Site Fidelity; Pathology Confirmation", "Common Misinterpretations": "Treating processing delay control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Processing Delay Control", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of processing delay control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses processing delay control using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in processing delay control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
83
575df935e8ca0edd1975c05772eaa99f38b63f196b0ee598cef40ebcb4554aaf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 207
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Population Applicability; Comparator Applicability; Outcome Applicability", "Common Misinterpretations": "Treating intervention applicability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Intervention Applicability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of intervention applicability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses intervention applicability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in intervention applicability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
207
d1cb3929ecc9b0976fba4823cb5a09ca13a634dd6efaa3ee43969eafe3250a95
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 364
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Immunotoxicity Evidence Strength; Species Extrapolation Validity; Human-Relevance of Toxicological Evidence", "Common Misinterpretations": "Treating toxicokinetic concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Toxicokinetic Concordance", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The degree of agreement in toxicokinetic across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses toxicokinetic concordance using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in toxicokinetic concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
364
6cda2f6cb167ccc21d05523345b2fef382432ceed6ebbcbe84fa383cddc4562e
Toxicological Mode-of-Action Support
The magnitude and credibility of independent evidence supporting toxicological mode-of-action.
bemo
BEMO:2000358
BEMO:2000358
365
05b74bda1c6ce0d8b99d4d0ffdc239f28a207c0631f25d1605f8d1438c74148a
8
Assesses toxicological mode-of-action support using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in toxicological mode-of-action support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating toxicological mode-of-action support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Calibration of Statistical Predictions
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 353
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Pharmacokinetic Adequacy; Bioavailability; Dose Proportionality", "Common Misinterpretations": "Treating pharmacodynamic adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Pharmacodynamic Adequacy", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The extent to which pharmacodynamic is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses pharmacodynamic adequacy using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in pharmacodynamic adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
353
99eef93672d2993b0ec131c09788c14463fd2a5166870e81b6b8d81c3637506b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 419
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Comparator Description Completeness; Recruitment Reporting Completeness; Participant Flow Completeness", "Common Misinterpretations": "Treating eligibility criteria completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Eligibility Criteria Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of eligibility criteria are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses eligibility criteria completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in eligibility criteria completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
419
02e23f40dd856c6f19f825d6191af4cb70f9ed2e604c47f5cbd7bf406e4d2381
Conflicting Interpretation Burden
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of conflicting interpretation burden.
bemo
BEMO:2000216
BEMO:2000216
223
54567aec6ae0bbdadfae015005c07328ab45d161c0cc6d36f9e5198150b2b29e
8
Assesses conflicting interpretation burden using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in conflicting interpretation burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating conflicting interpretation burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Cross-Platform Genomic Concordance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Dechallenge–Rechallenge Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Imputation Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Computation profile for Structural Identifiability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
On-Target Specificity
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of on-target specificity.
bemo
BEMO:2000018
BEMO:2000018
25
5b75004715c78ae57eb0dfbf1a6fe3b6212289ca6b06f39e35b36270d16b13e5
9
Assesses on-target specificity using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.
Material weakness in on-target specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating on-target specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Molecular, cellular, animal, translational, pharmacologic, and human studies
GRADE; FDA Biomarker; ClinGen; OHAT; OECD
https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Stable
Controlled BEMO LifecycleStatus value: Stable.
BEMO:4000024
Stable
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 494
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Temporal Precedence; Protocol Fidelity; Outcome Ascertainment Validity", "Common Misinterpretations": "Treating study design appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Study Design Appropriateness", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The extent to which study design is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses study design appropriateness using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in study design appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
494
7f9fb8e0c1ccb555a18bd7b6b07b3173d661a6707fab17a21ee632e59503b393
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 484
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Adherence Integrity; Intervention Classification Validity; Comparator Validity", "Common Misinterpretations": "Treating exposure classification validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Exposure Classification Validity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The degree to which exposure classification supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses exposure classification validity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in exposure classification validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
484
6d08552ad6c8b3c6d5a3c8ed35293fcfc624984197116fb158f6c3a04698789f
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 347
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Developmental Toxicity Evidence Strength; Toxicokinetic Concordance; Species Extrapolation Validity", "Common Misinterpretations": "Treating immunotoxicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Immunotoxicity Evidence Strength", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The magnitude and credibility of independent evidence supporting immunotoxicity evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses immunotoxicity evidence strength using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in immunotoxicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
347
577a4370c7a1dc270882fc655889fdc7a0b8ee20d5e6f81f1f733233db939fbd
Computation profile for Species Extrapolation Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 136
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Observed-to-Expected Ratio; Prognostic Transportability; Time-Dependent Discrimination", "Common Misinterpretations": "Treating prognostic added value as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Prognostic Added Value", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prognostic added value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses prognostic added value using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in prognostic added value can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
136
6a8e54c8d71bbaa2bc8361d6fc19eb1f50b8f4acd2b6d09e2b648f2337fa9298
Computation profile for Strand Bias
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 296
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Cutoff Validity; Method Comparison Agreement; Commutability", "Common Misinterpretations": "Treating measurement uncertainty as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Measurement Uncertainty", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of measurement uncertainty.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses measurement uncertainty using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in measurement uncertainty can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
296
02b8d5d020f47fcd1190b0fa52a15a86a5dd04748bbf2908e8c4e87ccdb4670f
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 288
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Carryover; Calibration Traceability; Reference Interval Validity", "Common Misinterpretations": "Treating hook effect risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Hook Effect Risk", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The probability or degree that hook effect introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses hook effect risk using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in hook effect risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
288
3d6a1547d736b7a958bf1c94a2df8cb5599400f10305a91c657ae1dce184650b
Computation profile for Type II Error Risk
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
has output specification
Connects a computation specification to its output specification.
BEMO:3000006
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 242
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Gene–Disease Validity; Population Frequency Compatibility; Segregation Evidence Strength", "Common Misinterpretations": "Treating variant pathogenicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Variant Pathogenicity Evidence Strength", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The magnitude and credibility of independent evidence supporting variant pathogenicity evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses variant pathogenicity evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in variant pathogenicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
242
ce239eedae0622eb1018ab2bfc52cd2d32ef0a5f6a874008c2d8838deb99bd4a
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 215
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Spectrum Representativeness; Geographic Consistency; Temporal Generalizability", "Common Misinterpretations": "Treating subgroup consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Subgroup Consistency", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "The degree of agreement in subgroup across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses subgroup consistency using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in subgroup consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
215
26cabc1a909ce4ef7d5d54dc609f85f9fea3829985a8f8573b3da07a3e225f5b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 93
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Confounding Risk; Residual Confounding Risk", "Common Misinterpretations": "Treating causal identifiability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Causal Identifiability", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of causal identifiability.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses causal identifiability using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in causal identifiability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
93
a6559bfe5ff78dd7455ed538b35cec6508c7c5507441bc614a71e2140188258e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 81
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Anatomical Site Fidelity; Tumor Purity; Cellularity Adequacy", "Common Misinterpretations": "Treating pathology confirmation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Pathology Confirmation", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of pathology confirmation.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses pathology confirmation using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in pathology confirmation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
81
e470a75855d7bb24a7f3957922ccdf78961ef4c923b6f754c0006ebb46c0367e
Computation profile for Geographic Consistency
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
has computation readiness status
Connects a computation specification to its readiness status.
BEMO:3000009
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 398
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Reanalysis Concordance; Protocol Reproducibility; Material Availability", "Common Misinterpretations": "Treating data provenance completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Data Provenance Completeness", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The extent to which all scientifically necessary components of data provenance are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses data provenance completeness using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in data provenance completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
398
da25c768c2ee710f0cf6b50b686a1880c8166dd651b4f062c14e952fc1bdf622
Computation profile for Effect-Modification Credibility
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 357
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Drug–Drug Interaction Evidence; Selectivity Profile; Potency Reproducibility", "Common Misinterpretations": "Treating receptor occupancy evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Receptor Occupancy Evidence", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of receptor occupancy evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses receptor occupancy evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in receptor occupancy evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
357
5025e34702ce03c8ed7e7aa3ed95dd9bde6bac1786183651ce608b34ffcc6a28
Computation profile for Area Under the Receiver Operating Characteristic Curve
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 405
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Protocol Reproducibility; Code Availability; Data Availability", "Common Misinterpretations": "Treating material availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Material Availability", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of material availability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses material availability using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in material availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
405
ad24de4b8dd0c9ae8c1a3d7d5fcdd166b9bd923967cbf8032882e26fce35feb2
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 431
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Raw-Data Availability; Quality-Control Reporting Completeness; Negative-Result Reporting", "Common Misinterpretations": "Treating processed-data availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Processed-Data Availability", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of processed-data availability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses processed-data availability using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in processed-data availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
431
3d4a65e6d3c66bcaca4ff40b6df33feaedf2f77c58697c32f80c96a00d211e95
Practical Identifiability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of practical identifiability.
bemo
BEMO:2000325
BEMO:2000325
332
cb110edeb57b36960477192ea5d1484823905d046fc0f22c33bc8eb340c3c4c3
8
Assesses practical identifiability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.
Material weakness in practical identifiability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating practical identifiability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Dataset / model / pathway / network / evidence body
Integrated omics, networks, pathways, mechanistic and dynamic systems models
Gene Ontology; Reactome; UniProt; GA4GH
https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
methods of assessment
Preserves the source methods of assessment.
BEMO:3200005
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 229
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Locus Heterogeneity Assessment; Variant Classification Stability; Conflicting Interpretation Burden", "Common Misinterpretations": "Treating genotype–phenotype concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Genotype–Phenotype Concordance", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The degree of agreement in genotype–phenotype across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses genotype–phenotype concordance using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in genotype–phenotype concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
229
baeee8e900a9ffef8f3ddd054b3a12f2233d89fc626f852fd31efc4b84b81685
Computation profile for Network Context Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 179
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Randomization in Experimental Allocation; Animal Model Face Validity; Animal Model Construct Validity", "Common Misinterpretations": "Treating blinding in experimental assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Blinding in Experimental Assessment", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of blinding in experimental assessment.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses blinding in experimental assessment using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in blinding in experimental assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
179
2a7451a420224db55e7f75f4a33818fc20cbc59d2bbc81f21fd652f7ea1a7931
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 248
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Single-Cell Feature Detection Rate; Spatial Transcriptomic Registration Accuracy; Cross-Platform Genomic Concordance", "Common Misinterpretations": "Treating cell-type annotation confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Cell-Type Annotation Confidence", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "The justified degree of certainty assigned to cell-type annotation given the quantity, quality, consistency, and limitations of supporting evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses cell-type annotation confidence using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cell-type annotation confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
248
f7aebc612affb172059739df32608fb7c4d69384f1ff8c565432595eea6b9937
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 489
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Carryover Effect Risk; Cluster Recruitment Bias Risk; Early Stopping Bias Risk", "Common Misinterpretations": "Treating period effect risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Period Effect Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that period effect introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses period effect risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in period effect risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
489
78e4475c3f570e4a3e966d2bc7b940517b5d7eaa514d888dc22a6ea5b68d38fb
Computation profile for Latent-Factor Stability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Penetrance Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of penetrance evidence.
bemo
BEMO:2000227
BEMO:2000227
234
83c8ddff797d0831554af5cae66094d559ff9fcf3cefc498edc05f2c18b6ee97
8
Assesses penetrance evidence using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.
Material weakness in penetrance evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.
Treating penetrance evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Mendelian disease, cancer genetics, association, segregation, and functional studies
ClinGen; ACMG AMP; STREGA; Gene Ontology
https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 374
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Mass Accuracy; Fragmentation Spectrum Quality; Post-Translational Modification Localization Confidence", "Common Misinterpretations": "Treating isotope pattern fidelity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Isotope Pattern Fidelity", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of isotope pattern fidelity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses isotope pattern fidelity using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in isotope pattern fidelity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
374
beec5d62727ea3baba10451d532801205b190af60b03027284c15989d51ba1e4
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 61
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Monitoring Biomarker Validity; Susceptibility/Risk Biomarker Validity; Response Biomarker Validity", "Common Misinterpretations": "Treating safety biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Safety Biomarker Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which safety biomarker supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses safety biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in safety biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
61
a6b6c6cf8ada87a1112583df9bc78595f4f0975ec0e8bbfc01ca6f00c2c0adf1
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 349
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Time–Concentration Profile Adequacy; Drug–Drug Interaction Evidence; Receptor Occupancy Evidence", "Common Misinterpretations": "Treating metabolite coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Metabolite Coverage", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The proportion and representativeness of the relevant metabolite captured by the evidence or measurement process.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses metabolite coverage using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in metabolite coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
349
4bb05358971c228c81cd2238f505b14da043cabfc3f68e3d2dd07884c7f72bc5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 301
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Sample Stability; Instrument Drift; Operator Variability", "Common Misinterpretations": "Treating reagent lot consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Reagent Lot Consistency", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The degree of agreement in reagent lot across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses reagent lot consistency using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reagent lot consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
301
66143063e8139ffa9de4cc6fb03a0eb1778bd92be05c901387fc7af071b0e433
source scientific definition
Preserves the exact scientific definition from the source catalog.
BEMO:3200001
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 205
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Transportability; Setting Applicability; Population Applicability", "Common Misinterpretations": "Treating generalizability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Generalizability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of generalizability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses generalizability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in generalizability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
205
27f8a61447192cbdeaaa7e84078183a3604b2e6ac6c0df13be6cfe9e4a09224a
generated by computation process
Connects a metric assessment to the process that generated it.
BEMO:3000020
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 154
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Completeness; Evidence Freshness; Evidence Robustness", "Common Misinterpretations": "Treating evidence coverage as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Coverage", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The proportion and representativeness of the relevant evidence captured by the evidence or measurement process.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses evidence coverage using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence coverage can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
154
a64cee9a2114ea7d5580ea8d59eaef0691fce1d63dd3e4fd9f3d63e0dac86a3d
Positive-Control Performance
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of positive-control performance.
bemo
BEMO:2000183
BEMO:2000183
190
98f43a472c5bc9ced611ad7ff89ac485c18a5c20ba0a8428164700738b649f75
8
Assesses positive-control performance using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in positive-control performance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating positive-control performance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 148
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Prediction Interval Adequacy; Information Size Adequacy; Multiplicity-Adjusted Credibility", "Common Misinterpretations": "Treating cumulative evidence stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Cumulative Evidence Stability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cumulative evidence stability.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses cumulative evidence stability using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cumulative evidence stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
148
e87943e84aae44a1eeb6300f5a52416f0528fe0f3bf76f40f0a2a4486c87ea6a
Computation profile for Comparator Description Completeness
0.1.0
value = numerator / denominator, with denominator > 0
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"divide","arguments":["numerator","denominator"],"constraints":["denominator > 0"]}
numerator; denominator; operational_definition; assessment_context
weight; stratum; confidence_level
xsd:decimal
Proportion or percentage (0–1 or 0–100%)
0.0
1.0
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; required when interpreted probabilistically.
Source maturity: Established; BEMO computation profile requires independent validation.
Biomarker Reliability
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biomarker reliability.
bemo
BEMO:2000030
BEMO:2000030
37
74c0c7f33ecc210cbe2fa1dd5ddc15c31273005f3e5caaa4afc370a3f6c94a4e
8
Assesses biomarker reliability using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.
Material weakness in biomarker reliability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating biomarker reliability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Biomarker development, qualification, endpoint and surrogate validation studies
FDA Biomarker; BEST; EMA E16; REMARK
https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
computation profile IRI
Links a metric class to its profile using an annotation-safe IRI.
BEMO:3200012
Computation profile for Gene–Disease Validity
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Mature; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 55
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Pharmacodynamic Biomarker Validity; Safety Biomarker Validity; Susceptibility/Risk Biomarker Validity", "Common Misinterpretations": "Treating monitoring biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Monitoring Biomarker Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which monitoring biomarker supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses monitoring biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in monitoring biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
55
cc605972d7119b4dabf5e0bc927d29fa05979e706703654bb3ea93c2ebe95461
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 98
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Dose–Response Support; Mediation Evidence Strength; Effect-Modification Credibility", "Common Misinterpretations": "Treating dechallenge–rechallenge support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Dechallenge–Rechallenge Support", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The magnitude and credibility of independent evidence supporting dechallenge–rechallenge.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses dechallenge–rechallenge support using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in dechallenge–rechallenge support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
98
9e166f37b12b35c895b18fc7e56e04526e22004c0b70dee562c1cf80c5d17e69
Computation profile for Random-Seed Stability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
computation readiness status
A controlled concept describing how close a metric is to executable, validated computation.
bemo
BEMO:0000402
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 395
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Analytical Reproducibility; Experimental Reproducibility; Interlaboratory Reproducibility", "Common Misinterpretations": "Treating computational reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Computational Reproducibility", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The degree to which computational reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses computational reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in computational reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
395
c15f070192e578e50deb91384b3e8099e5c4b1a0c28cc84d1ac6a3163d36cd17
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 72
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "RNA Integrity; Protein Integrity; Microbial Contamination", "Common Misinterpretations": "Treating dna integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "DNA Integrity", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of dna integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses dna integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in dna integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
72
494036e6478a34bfe608bc3b1267e04a777d16e96430109ab2ba9b01e99c6693
Icterus Interference
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of icterus interference.
bemo
BEMO:2000069
BEMO:2000069
76
d0f3346c9caa8abb9b49bded0ab6e26e5f13d5a8fb5a20971f64067dd82054ea
8
Assesses icterus interference using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.
Material weakness in icterus interference can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.
Treating icterus interference as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
All studies using human or animal biospecimens
BRISQ; ISO 15189; REMARK
https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for No-Observed-Adverse-Effect Level Robustness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
composite of metric
Connects a composite metric to a component metric.
BEMO:3000018
Quantitative Proportion Computation
Controlled BEMO ComputationMode value: QuantitativeProportionComputation.
BEMO:4000008
QuantitativeProportionComputation
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 235
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Founder-Effect Assessment; Locus Heterogeneity Assessment; Genotype–Phenotype Concordance", "Common Misinterpretations": "Treating phenocopy risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Phenocopy Risk", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The probability or degree that phenocopy introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses phenocopy risk using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in phenocopy risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
235
eba377fe2c190b6390cd804f64874ad58540debe287184cfa259e70d97d1fcc8
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 97
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "No-Interference Plausibility; Causal Contrast Clarity; Target Trial Emulation Fidelity", "Common Misinterpretations": "Treating correct temporal ordering as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Correct Temporal Ordering", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of correct temporal ordering.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses correct temporal ordering using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in correct temporal ordering can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
97
b11d505f3985b03d1ab78fe50a79558d187aeacc3fabbc39b5dc3da61a0cfbe2
Missing Evidence Risk
The probability or degree that missing evidence introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000158
BEMO:2000158
165
ffec67315125f9639a3601d51e352906c34b1c9847a274942f934888c1188977
8
Assesses missing evidence risk using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.
Material weakness in missing evidence risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating missing evidence risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Body of evidence / synthesis / outcome
Systematic reviews, meta-analyses, evidence profiles, guidelines
GRADE; PRISMA; AMSTAR 2; RoB
https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Differential Verification Bias Risk
The probability or degree that differential verification bias introduces systematic distortion into a biomedical estimate or conclusion.
bemo
BEMO:2000118
BEMO:2000118
125
2805e0d0838d99254f4675a5517e72c35cb3e1cb875fbc214edb03376568c858
10
Assesses differential verification bias risk using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.
Material weakness in differential verification bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating differential verification bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies
QUADAS-2; STARD; TRIPOD; REMARK
https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Mature
Computation profile for Systems-Level Emergence Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
has formula status
Connects a computation specification to formula availability status.
BEMO:3000010
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 379
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Dynamic Range Coverage; Ion Suppression Assessment; Retention-Time Stability", "Common Misinterpretations": "Treating missing-value burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Missing-Value Burden", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of missing-value burden.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses missing-value burden using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in missing-value burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
379
770bd2fa8c1eedf14831e8e8caca6d53a1ad4a7b7272433891be0eac41a4e30c
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 41
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Biomarker-Outcome Association Strength; Biomarker Clinical Relevance", "Common Misinterpretations": "Treating clinical validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Clinical Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which clinical supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses clinical validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in clinical validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
41
628005d99d3e18a7b100ce400c64fc5218e46beeaa2cfec80dbac2445270cead
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 378
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Proteoform Identification Confidence; Metabolite Annotation Level; Spectral Library Match Quality", "Common Misinterpretations": "Treating metabolite identification confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Metabolite Identification Confidence", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The justified degree of certainty assigned to metabolite identification given the quantity, quality, consistency, and limitations of supporting evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses metabolite identification confidence using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in metabolite identification confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
378
dd08774f62d4c8036bc0f7c590a94eca58c7ca0613c38dec0e386cab59cc0dbd
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 360
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Receptor Occupancy Evidence; Potency Reproducibility; Efficacy Reproducibility", "Common Misinterpretations": "Treating selectivity profile as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Selectivity Profile", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of selectivity profile.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses selectivity profile using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in selectivity profile can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
360
2957579a3d734fc808c8a89eb160fb182b5f5a0f68ba40c4eecc1921dea93153
Genomics and Transcriptomics metric
Category of biomedical evidence metrics concerned with genomics and transcriptomics.
bemo
BEMO:1100010
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 407
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Researcher-Degrees-of-Freedom Sensitivity; Specification-Curve Robustness", "Common Misinterpretations": "Treating multiverse analysis robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Specialized / infrequent", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Developing", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Multiverse Analysis Robustness", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of multiverse analysis robustness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses multiverse analysis robustness using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in multiverse analysis robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
407
80472d064c943a6566bee4bd605245ecac370ebde0eb5ae577d7687482d47cc8
Context Dependent Direction
Controlled BEMO Directionality value: ContextDependentDirection.
BEMO:4000021
ContextDependentDirection
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 203
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Context Sensitivity; Real-World Evidence Alignment", "Common Misinterpretations": "Treating ecological validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Ecological Validity", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "The degree to which ecological supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses ecological validity using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in ecological validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
203
04e73707dc619457f6718e0be03029ef6e99208f68cbcad30d9f5bb20b2cd94b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 414
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Data-Sharing Transparency; Materials-and-Reagents Reporting Completeness; Metadata Completeness", "Common Misinterpretations": "Treating code-sharing transparency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Code-Sharing Transparency", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of code-sharing transparency.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses code-sharing transparency using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in code-sharing transparency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
414
d85d4b116984de9e36c9410fc248aae4b536153c50f4651cf7e48c6b558b49eb
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 281
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Hook Effect Risk; Reference Interval Validity; Cutoff Validity", "Common Misinterpretations": "Treating calibration traceability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Calibration Traceability", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of calibration traceability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses calibration traceability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in calibration traceability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
281
7d6efe540a48fdb6c1d86b381d867ad016fd03d3ba3039e34ba0dd8d59bcbf8d
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 173
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Sex as a Biological Variable Adequacy; Genetic Background Control; Housing and Husbandry Control", "Common Misinterpretations": "Treating age appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Age Appropriateness", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The extent to which age is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses age appropriateness using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in age appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
173
6f50922de3489bf44a56c621fcc8933973821a524b4c3ba13cb746276b4a9bea
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 112
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Immortal-Time Bias Risk; Selection-on-Survival Bias Risk; Quantitative Bias Analysis Robustness", "Common Misinterpretations": "Treating reverse-causation risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Reverse-Causation Risk", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The probability or degree that reverse-causation introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses reverse-causation risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reverse-causation risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
112
17fc4b74f3434a4ae22f908b09530b9a1be33f82af499c89239235d78d7e1375
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 109
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Exchangeability Plausibility; Consistency Assumption Plausibility; No-Interference Plausibility", "Common Misinterpretations": "Treating positivity adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Positivity Adequacy", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The extent to which positivity is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses positivity adequacy using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in positivity adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
109
be2ac5c0345d954f0ccc1ce41f05b792425f3fb9c158083a6db1928102b69608
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 110
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Selection-on-Survival Bias Risk; E-value Strength; Assumption Sensitivity", "Common Misinterpretations": "Treating quantitative bias analysis robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Quantitative Bias Analysis Robustness", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of quantitative bias analysis robustness.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses quantitative bias analysis robustness using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in quantitative bias analysis robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
110
7ee1bd197cfd8211ed7ac785b0e49ec8e73324814e415213789e479477f73c74
Pharmacology and Toxicology metric
Metrics assessing pharmacologic activity, exposure, efficacy, selectivity, and toxicologic evidence.
bemo
BEMO:1000008
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 239
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Population Frequency Compatibility; De Novo Evidence Strength; Allelic Evidence Strength", "Common Misinterpretations": "Treating segregation evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Segregation Evidence Strength", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The magnitude and credibility of independent evidence supporting segregation evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses segregation evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in segregation evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
239
4f9bdcfda63260ced3f40796d43228eea03ead68f517bc87939698d4f5d4faf2
machine formula language
Language used for a machine-readable formula.
BEMO:3100004
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 196
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Biological Replicate Adequacy; Sample Size Justification; Randomization in Experimental Allocation", "Common Misinterpretations": "Treating technical replicate adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Technical Replicate Adequacy", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The extent to which technical replicate is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses technical replicate adequacy using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in technical replicate adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
196
8f5bd5eed372632eafdff28f1c8035067d4634c207012fb089ff55a67d705176
Computation profile for Method Comparison Agreement
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Bayes Factor Evidence
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of bayes factor evidence.
bemo
BEMO:2000432
BEMO:2000432
439
ea8be0499c9c5441e47cac85f3e884bcc28099d593610c01dc6ad24ba5c3d0f3
8
Assesses bayes factor evidence using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in bayes factor evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating bayes factor evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 321
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Network Node Confidence; Pathway Enrichment Consistency; Pathway Topology Support", "Common Misinterpretations": "Treating module stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Module Stability", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of module stability.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses module stability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in module stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
321
191acd0c1abb112365dc2e82b9a1c639642c0798eb52983d078ac92e3f0fbeec
Animal Model Face Validity
The degree to which animal model face supports the intended scientific interpretation without material systematic error.
bemo
BEMO:2000168
BEMO:2000168
175
1bc9d94635678f728d2a96fd7016f3d61e24c78d9e33c5093213c9d715511534
10
Assesses animal model face validity using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in animal model face validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating animal model face validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Mature
Computation profile for Hemolysis Burden
0.1.0
Apply a versioned risk rubric or calibrated probability model defined for the metric and context.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_risk_category","probability","percentage"]}
risk_factors; assessment_context; rubric_or_model_version; evidence_records
weights; thresholds; calibration_dataset
xsd:string_or_decimal
Ordinal risk rating or quantitative percentage/probability; lower is generally better
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for probability outputs; rubric reliability required for ordinal outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
validated against
Connects a metric profile or assessment to a benchmark or reference standard.
BEMO:3000019
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 478
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Intervention Classification Validity; Control Group Appropriateness; Temporal Precedence", "Common Misinterpretations": "Treating comparator validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Comparator Validity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The degree to which comparator supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses comparator validity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in comparator validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
478
de06a641ed9f64a7c11259e9a1d1346140c2469d57bbd106c6257d947909a5a3
Intralaboratory Repeatability
The degree to which intralaboratory repeatability yields concordant results under the specified repeated-analysis or repeated-measurement conditions.
bemo
BEMO:2000397
BEMO:2000397
404
44bfb4ae3b82bb963a97cc8f3dfed3bd610b7903405993c8f3031931f746d1ce
8
Assesses intralaboratory repeatability using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.
Material weakness in intralaboratory repeatability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating intralaboratory repeatability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All experimental, computational, clinical, and omics studies
PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0
https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Ordinal Or Normalized Score Scale
Controlled BEMO ScaleType value: OrdinalOrNormalizedScoreScale.
BEMO:4000002
OrdinalOrNormalizedScoreScale
Blinding in Experimental Assessment
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of blinding in experimental assessment.
bemo
BEMO:2000172
BEMO:2000172
179
2a7451a420224db55e7f75f4a33818fc20cbc59d2bbc81f21fd652f7ea1a7931
8
Assesses blinding in experimental assessment using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in blinding in experimental assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating blinding in experimental assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 352
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Adverse Outcome Pathway Support; Genotoxicity Evidence Strength; Carcinogenicity Evidence Strength", "Common Misinterpretations": "Treating organ-specific toxicity evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Organ-Specific Toxicity Evidence", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of organ-specific toxicity evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses organ-specific toxicity evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in organ-specific toxicity evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
352
277a918d22b52db9e79eabf5451f64320f342f326be1cc2382f16c69ec72f0a9
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 257
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Genome Coverage Uniformity; Read Quality; Duplicate Read Burden", "Common Misinterpretations": "Treating mapping quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Mapping Quality", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of mapping quality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Read- and variant-level quality-control summaries; replicate concordance; orthogonal confirmation; benchmarking against reference materials.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses mapping quality using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in mapping quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
257
aef9eeffa971170a3db3642e806206bdf35806d1d8cc427e16ad9106f7e0207c
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 144
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Prognostic Transportability; Competing-Risk Model Validity", "Common Misinterpretations": "Treating time-dependent discrimination as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Time-Dependent Discrimination", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of time-dependent discrimination.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses time-dependent discrimination using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in time-dependent discrimination can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
144
5232bfe9f1dc1ddd88662ddb65a550d0d1fc5338bb4a530ed1072ed6078116bf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 466
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Confidence Interval Compatibility; Type I Error Control; Type II Error Risk", "Common Misinterpretations": "Treating statistical power as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Statistical Power", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of statistical power.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses statistical power using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in statistical power can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
466
b5866270dc0b1ae7d14dc9e42f0ff6a1d28896f404d73ddeb44c5f7fdaf6b51a
threshold policy
Policy governing thresholds.
BEMO:3100019
Analytical Measurement Range
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of analytical measurement range.
bemo
BEMO:2000268
BEMO:2000268
275
40562e2bdf0f0c373312aa0a75941fd6a83f056ba5201b300a46242b3f5457df
10
Assesses analytical measurement range using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.
Material weakness in analytical measurement range can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating analytical measurement range as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Specimen / assay / run / laboratory / study
Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies
FDA Biomarker; CLSI; ISO 15189; MIQE
https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common within specialty
Established
Computation profile for Single-Cell Feature Detection Rate
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 44
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Biomarker Clinical Relevance; Biomarker Qualification Strength; Biomarker Reliability", "Common Misinterpretations": "Treating context-of-use validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Context-of-Use Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which context-of-use supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses context-of-use validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in context-of-use validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
44
170c4bee423f9901b11abd5c6c7203f1253da40dbd46a33c459bac07471058e8
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 375
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Retention-Time Stability; Isotope Pattern Fidelity; Fragmentation Spectrum Quality", "Common Misinterpretations": "Treating mass accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Mass Accuracy", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The closeness of mass to the accepted reference or true value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses mass accuracy using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in mass accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
375
ebb0e828e6a056d20bce6788aee10b743cfcc2fc175a82fd693585711880ff73
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 283
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Method Comparison Agreement; Sample Stability; Reagent Lot Consistency", "Common Misinterpretations": "Treating commutability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Commutability", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of commutability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses commutability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in commutability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
283
0d8876821c8ce33e4a6f72fcdd334c759d21fb2784bd7f9f779179758d7849c8
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 317
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Pathway Topology Support; Knowledge-Graph Provenance Quality; Entity Resolution Accuracy", "Common Misinterpretations": "Treating knowledge-graph evidence completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Developing", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Knowledge-Graph Evidence Completeness", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The extent to which all scientifically necessary components of knowledge-graph evidence are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses knowledge-graph evidence completeness using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in knowledge-graph evidence completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
317
082de1cac908d5d8b791ddee2974743d0e0eae71997e72080a5d978e931f8b6c
has optional input specification
Connects a computation specification to an optional input specification.
BEMO:3000005
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 458
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Type II Error Risk; Model Specification Adequacy; Model Fit", "Common Misinterpretations": "Treating multiplicity control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Multiplicity Control", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of multiplicity control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses multiplicity control using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in multiplicity control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
458
282fd29ea793d9f6346d4b34598e7ca30406dad2a72b3f3167d354ca92803ec0
closely related metric
Source-declared closely related metric.
BEMO:3200013
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 90
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Collection Procedure Consistency; Cold Ischemia Control; Time-to-Fixation Adequacy", "Common Misinterpretations": "Treating warm ischemia control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Warm Ischemia Control", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of warm ischemia control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses warm ischemia control using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in warm ischemia control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
90
293c5ebda3d3695852f06644549a5b920266b83358bf49b383b5d90078fcdde6
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 491
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Study Design Appropriateness; Outcome Ascertainment Validity; Follow-up Completeness", "Common Misinterpretations": "Treating protocol fidelity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Protocol Fidelity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of protocol fidelity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses protocol fidelity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in protocol fidelity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
491
ae404e8a9df9a665347ede4db7ae7370bdc6c77c15da3474e23c516ad81dd25e
Computation profile for Intralaboratory Repeatability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 94
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Effect-Modification Credibility; Time-Varying Confounding Control; Immortal-Time Bias Risk", "Common Misinterpretations": "Treating collider bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Collider Bias Risk", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The probability or degree that collider bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses collider bias risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in collider bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
94
0d0ee23715a02b538d00c3772dc0b2a1a64aba9fb6a21ecd62845d74c88b63ca
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 246
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Hardy–Weinberg Equilibrium Compatibility; Population Stratification Control; Relatedness Control", "Common Misinterpretations": "Treating batch-effect control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Batch-Effect Control", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of batch-effect control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses batch-effect control using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in batch-effect control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
246
0751d3491af658db1257b3b04a366a751c7fd0fdfb5e615ebbae3dab6970c7c9
metric input specification
An information content entity specifying a required or optional input role for a metric computation.
bemo
BEMO:0000101
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 135
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Diagnostic Specificity; Negative Predictive Value; Positive Likelihood Ratio", "Common Misinterpretations": "Treating positive predictive value as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Positive Predictive Value", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of positive predictive value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses positive predictive value using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in positive predictive value can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
135
89034784369de0142956f172eb3ca87ccad9b4cc69ac8a45b8a92adf9cc1731f
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 412
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Method Reproducibility; Inferential Reproducibility; Reanalysis Concordance", "Common Misinterpretations": "Treating result reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Result Reproducibility", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The degree to which result reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses result reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in result reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
412
60121a2998efa943018d3ff7dd4130d4332ee279da0e8111b650ea175e9dacce
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 35
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Biomarker-Outcome Association Strength; Context-of-Use Validity; Biomarker Qualification Strength", "Common Misinterpretations": "Treating biomarker clinical relevance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Biomarker Clinical Relevance", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biomarker clinical relevance.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses biomarker clinical relevance using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biomarker clinical relevance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
35
70bbbca5f05350dbfed193659c18ff0eef66a6552528618a931911c0bbfe7878
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 70
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Warm Ischemia Control; Time-to-Fixation Adequacy; Fixation Adequacy", "Common Misinterpretations": "Treating cold ischemia control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Cold Ischemia Control", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cold ischemia control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses cold ischemia control using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cold ischemia control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
70
450f9e94b009036abe7a13fe3055116e7c4959857a7291ed7e337a746d7bf722
Computation profile for Mechanistic Causality Strength
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 380
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Metabolic Feature Reproducibility; Cross-Omics Concordance", "Common Misinterpretations": "Treating pathway enrichment robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Pathway Enrichment Robustness", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of pathway enrichment robustness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses pathway enrichment robustness using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in pathway enrichment robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
380
1a891a794290e0f5eee9ca88b64bc6b46d5840f63c8709e0e546399158fbb914
uncertainty method
Acceptable uncertainty estimation method.
BEMO:3100017
Experimental Biology and Animal Research metric
Category of biomedical evidence metrics concerned with experimental biology and animal research.
bemo
BEMO:1100007
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 480
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Comparator Validity; Temporal Precedence; Study Design Appropriateness", "Common Misinterpretations": "Treating control group appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Control Group Appropriateness", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The extent to which control group is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses control group appropriateness using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in control group appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
480
47dcb7d2162b3c8496d326c0cd9b9dfbef0e0b4032c2c7d93f31f67139c5990f
Rejected
Controlled BEMO ApprovalStatus value: Rejected.
BEMO:4000030
Rejected
Computation profile for Model Specification Adequacy
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
computation mode
A controlled concept describing the general computational approach used by a metric.
bemo
BEMO:0000401
Candidate
Study Design and Internal Validity metric
Category of biomedical evidence metrics concerned with study design and internal validity.
bemo
BEMO:1100018
Candidate
Risk Scale
Controlled BEMO ScaleType value: RiskScale.
BEMO:4000006
RiskScale
Obsolete
Controlled BEMO LifecycleStatus value: Obsolete.
BEMO:4000026
Obsolete
Computation profile for Outlier Influence Robustness
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 211
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Ecological Validity", "Common Misinterpretations": "Treating real-world evidence alignment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Real-World Evidence Alignment", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of real-world evidence alignment.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses real-world evidence alignment using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in real-world evidence alignment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
211
2bcc973aa3f12efded260e2aef1541636262ee110856a34255ce0823d91dfb20
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 333
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Entity Resolution Accuracy; Ontology Annotation Completeness; Ontology Evidence-Code Strength", "Common Misinterpretations": "Treating relation evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Relation Evidence Strength", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The magnitude and credibility of independent evidence supporting relation evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses relation evidence strength using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in relation evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
333
9009911ac3f6266273d49aa664b528314dd25468ab2969a5c3dfb8593502a99c
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 37
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Biomarker Qualification Strength; Biomarker Responsiveness; Biomarker Sensitivity to Change", "Common Misinterpretations": "Treating biomarker reliability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Biomarker Reliability", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biomarker reliability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses biomarker reliability using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biomarker reliability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
37
74c0c7f33ecc210cbe2fa1dd5ddc15c31273005f3e5caaa4afc370a3f6c94a4e
Physical Measurement Scale
Controlled BEMO ScaleType value: PhysicalMeasurementScale.
BEMO:4000003
PhysicalMeasurementScale
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 51
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Trial-Level Surrogacy; Outcome Relevance; Endpoint Reliability", "Common Misinterpretations": "Treating endpoint validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Endpoint Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which endpoint supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses endpoint validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in endpoint validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
51
f79c486fa1f56b17619593da318c44a3704cbd1294be41ae30a2b38ceb9410d1
source study types
Preserves the source applicable study types.
BEMO:3200009
metric applicability profile
An information content entity specifying evidence levels, study types, domains, and contexts to which a metric applies.
bemo
BEMO:0000103
Candidate
source references
Preserves the source reference list.
BEMO:3200011
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 26
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Mechanistic Causality Strength; Cell-Type Specificity; Tissue Specificity", "Common Misinterpretations": "Treating pathway-level support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Pathway-Level Support", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The magnitude and credibility of independent evidence supporting pathway-level.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses pathway-level support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in pathway-level support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
26
a1a0e0a377d9f949a45a0ad0a040b0443e29f24bfe701f3b015420b52cd21ab2
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 439
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Fragility Index; Posterior Probability Strength; Prior Sensitivity", "Common Misinterpretations": "Treating bayes factor evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Bayes Factor Evidence", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of bayes factor evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses bayes factor evidence using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in bayes factor evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
439
ea8be0499c9c5441e47cac85f3e884bcc28099d593610c01dc6ad24ba5c3d0f3
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 182
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Exclusion-Criteria Prespecification; Positive-Control Performance; Negative-Control Performance", "Common Misinterpretations": "Treating experimental batch randomization as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Experimental Batch Randomization", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of experimental batch randomization.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses experimental batch randomization using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in experimental batch randomization can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
182
1291027c6f53e5902a86a2cf22b3cca02564dbc940590bb8b8bcfd5004b998b7
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 455
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Missing-Data Mechanism Plausibility; Imputation Validity; Outlier Influence Robustness", "Common Misinterpretations": "Treating missing-data sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Missing-Data Sensitivity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of missing-data sensitivity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses missing-data sensitivity using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in missing-data sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
455
6708217d617153345c000ff0b34cc96bd34facf9514e5d1a84b3e56623ffb545
has scale type
Connects a computation specification to a controlled scale type.
BEMO:3000007
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 234
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Co-segregation Likelihood; Expressivity Consistency; Founder-Effect Assessment", "Common Misinterpretations": "Treating penetrance evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Penetrance Evidence", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of penetrance evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses penetrance evidence using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in penetrance evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
234
83c8ddff797d0831554af5cae66094d559ff9fcf3cefc498edc05f2c18b6ee97
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 232
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Variant Phase Evidence; Hotspot/Functional-Domain Evidence; Null-Variant Quality", "Common Misinterpretations": "Treating loss-of-function mechanism validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Loss-of-Function Mechanism Validity", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The degree to which loss-of-function mechanism supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses loss-of-function mechanism validity using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in loss-of-function mechanism validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
232
a389d5411c414282451fe5495bf2c1c444f06ae12cd17457edf93df2953dc32e
External Validity, Biomarkers, and Clinical Evidence metric
Metrics assessing generalizability, diagnostic or prognostic evidence, biomarkers, endpoints, and clinical applicability.
bemo
BEMO:1000005
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 100
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Quantitative Bias Analysis Robustness; Assumption Sensitivity; Falsification Test Support", "Common Misinterpretations": "Treating e-value strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Specialized / infrequent", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Developing", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "E-value Strength", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The magnitude and credibility of independent evidence supporting e-value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ratio scale; null typically 1", "What It Measures": "Assesses e-value strength using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in e-value strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
100
8587ae49790e995db7d3bf84506f3bebdce4263532f6e5dcbd136229f5da442f
Technical Artifact Exclusion
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of technical artifact exclusion.
bemo
BEMO:2000188
BEMO:2000188
195
8d21017136fa388d87b531d9b552fc3f6787a0b2615866e0c0a1ca9ae2fb0bbf
8
Assesses technical artifact exclusion using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.
Material weakness in technical artifact exclusion can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating technical artifact exclusion as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
In vitro, ex vivo, organoid, animal, and preclinical experiments
ARRIVE 2.0; SYRCLE; OECD
https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Computation profile for Allelic Balance
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Computation profile for Pathway-Level Support
0.1.0
Apply a versioned, prespecified domain rubric or validated normalized scoring model.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED","allowed_outputs":["ordinal_category","normalized_score"]}
evidence_records; assessment_context; rubric_version; operational_definition
weights; thresholds; expert_adjudication
xsd:string_or_decimal
Ordinal rubric, domain judgment, or normalized score
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
false
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
Required for normalized scores; inter-rater reliability required for human rubrics.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 222
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Functional Variant Evidence; Phenotypic Specificity for Variant; Variant Phase Evidence", "Common Misinterpretations": "Treating computational variant evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Computational Variant Evidence", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of computational variant evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses computational variant evidence using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in computational variant evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
222
1f84faa3d8c87f3b1ed44b62077d2c489782ab72b149d6009c127395a0070a10
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 432
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Protocol Availability; Prespecified Analysis Adherence; Outcome Definition Completeness", "Common Misinterpretations": "Treating prospective registration as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Prospective Registration", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prospective registration.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses prospective registration using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in prospective registration can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
432
32ae2352b0f17339d30893a08c10c681c7fa3992e6b60a3343dae6e95aa8dda5
Distributional Assumption Adequacy
The extent to which distributional assumption is sufficient and fit for the stated biomedical inference.
bemo
BEMO:2000438
BEMO:2000438
445
79608c041c0babc1613cca804b9e22a2fcbadbeabfda7d7b0bf2336528db7db1
8
Assesses distributional assumption adequacy using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.
Material weakness in distributional assumption adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating distributional assumption adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
All quantitative biomedical studies
CONSORT; STROBE; TRIPOD; REMARK; ICH E9
https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Common
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 410
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Inferential Reproducibility; Data Provenance Completeness; Protocol Reproducibility", "Common Misinterpretations": "Treating reanalysis concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Reanalysis Concordance", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The degree of agreement in reanalysis across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses reanalysis concordance using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reanalysis concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
410
07f0b47b85e307e7887d3c938a89cbcad651d37d8bfa888ea4ce710d25410a29
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 449
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Clinical Relevance of Effect; Bayes Factor Evidence; Posterior Probability Strength", "Common Misinterpretations": "Treating fragility index as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Fragility Index", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of fragility index.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses fragility index using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in fragility index can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
449
045c39ed086cc31d598f1fe91547a90170ea8cc0b684f22d5338144e750a0843
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 60
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Susceptibility/Risk Biomarker Validity; Surrogate Endpoint Validity; Individual-Level Surrogacy", "Common Misinterpretations": "Treating response biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Response Biomarker Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which response biomarker supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses response biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in response biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
60
a8e8ef21a701fbb547c028557e4678bc79ebeda610eb6926bf0061b6e9a016c4
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 342
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Metabolite Coverage; Receptor Occupancy Evidence; Selectivity Profile", "Common Misinterpretations": "Treating drug–drug interaction evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Drug–Drug Interaction Evidence", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of drug–drug interaction evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses drug–drug interaction evidence using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in drug–drug interaction evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
342
3a82aa8a1751fa8372cd9735fd975781682a359be68aacfb1d0d78e896d96ce5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 319
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Cross-Layer Directional Concordance; Network Reconstruction Robustness; Network Edge Confidence", "Common Misinterpretations": "Treating latent-factor stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Latent-Factor Stability", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of latent-factor stability.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses latent-factor stability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in latent-factor stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
319
be1e726b2fd3735d5492ec8f375d473b36cbff860abc6d682d2065974c115c50
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 472
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Detection Bias Risk; Selective Outcome Reporting Risk; Protocol Deviation Risk", "Common Misinterpretations": "Treating attrition bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Attrition Bias Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that attrition bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses attrition bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in attrition bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
472
0783cbbb1a7caf7ef448d6830ae564e73010beb87e2e08710ddb494b6b76ccb3
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 199
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Intervention Applicability; Outcome Applicability; Spectrum Representativeness", "Common Misinterpretations": "Treating comparator applicability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Comparator Applicability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of comparator applicability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses comparator applicability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in comparator applicability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
199
607cceee390b7cb8ce577becf5db87c7c9a4643da6c8fbe5bd44948e1b9e3853
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 65
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Processing Delay Control; Pathology Confirmation; Tumor Purity", "Common Misinterpretations": "Treating anatomical site fidelity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Anatomical Site Fidelity", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of anatomical site fidelity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses anatomical site fidelity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in anatomical site fidelity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
65
8fa3b411fb892fff8cc31c078e25770eb74980a525cbc8e76d63bcb0f15bdaf8
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 223
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Variant Classification Stability", "Common Misinterpretations": "Treating conflicting interpretation burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Conflicting Interpretation Burden", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of conflicting interpretation burden.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses conflicting interpretation burden using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in conflicting interpretation burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
223
54567aec6ae0bbdadfae015005c07328ab45d161c0cc6d36f9e5198150b2b29e
has computation mode
Connects a computation specification to a computation-mode category.
BEMO:3000008
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 189
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Vehicle-Control Validity; Technical Artifact Exclusion", "Common Misinterpretations": "Treating orthogonal validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Orthogonal Validation", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of orthogonal validation.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses orthogonal validation using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in orthogonal validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
189
86d52a57d493edc61e94352ed7d2d8f4feaaa492b37b0afb047315bb8706c00e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 95
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Causal Identifiability; Residual Confounding Risk; Unmeasured Confounding Sensitivity", "Common Misinterpretations": "Treating confounding risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Confounding Risk", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The probability or degree that confounding introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses confounding risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in confounding risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
95
cebd58c93ed86b363870e131c5d23217522d15af0abb7eac8956ec81e8515565
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 231
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Phenocopy Risk; Genotype–Phenotype Concordance; Variant Classification Stability", "Common Misinterpretations": "Treating locus heterogeneity assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Locus Heterogeneity Assessment", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of locus heterogeneity assessment.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Forest plots; heterogeneity statistics; tau-squared; prediction intervals; funnel plots; regression or selection models; sensitivity analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses locus heterogeneity assessment using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in locus heterogeneity assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
231
bd3af304ec15826df5f67db7f079c0ba66a1f5f5a3e852d600550186267f50cb
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 225
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Penetrance Evidence; Founder-Effect Assessment; Phenocopy Risk", "Common Misinterpretations": "Treating expressivity consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Expressivity Consistency", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The degree of agreement in expressivity across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses expressivity consistency using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in expressivity consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
225
d100842d4d2c367ecd1122df00784975aa0a773a92bbb5c68a8ff23fbbd4dfc4
Biological Mechanism and Experimental Evidence metric
Metrics assessing biological plausibility, mechanistic evidence, experimental biology, and model-system evidence.
bemo
BEMO:1000006
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 101
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Mediation Evidence Strength; Collider Bias Risk; Time-Varying Confounding Control", "Common Misinterpretations": "Treating effect-modification credibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Effect-Modification Credibility", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of effect-modification credibility.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses effect-modification credibility using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in effect-modification credibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
101
f66521085db07e7bd17f7b93b07cf74e57b1c85461a2555c1146f672069098b2
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 227
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Case-Control Evidence Strength; Computational Variant Evidence; Phenotypic Specificity for Variant", "Common Misinterpretations": "Treating functional variant evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Functional Variant Evidence", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of functional variant evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses functional variant evidence using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in functional variant evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
227
7a71bec12dfb580d6d32593ad450868d51a755f2a8b35e7eee350ad98d24c73e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 74
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Storage Temperature Control; Transport Condition Integrity; Processing Delay Control", "Common Misinterpretations": "Treating freeze–thaw burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Freeze–Thaw Burden", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of freeze–thaw burden.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses freeze–thaw burden using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in freeze–thaw burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
74
40c5c886eacc67a9addeaa92c47d69bf9fbd4b1d2a0042a2efd3611fd5cebb4b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 163
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Robustness; Counterevidence Strength; Missing Evidence Risk", "Common Misinterpretations": "Treating evidence triangulation strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Triangulation Strength", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The magnitude and credibility of independent evidence supporting evidence triangulation.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses evidence triangulation strength using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence triangulation strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
163
b9d887ba211f3aea4699afd087a0aa12fb8ca924209ee34ccda2bfd968c02a4a
Mixture Interaction Assessment
A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of mixture interaction assessment.
bemo
BEMO:2000343
BEMO:2000343
350
241ffa4364ee9325fbb3832728344d8484375a1aca172f524b6f871b9dc07687
8
Assesses mixture interaction assessment using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.
Material weakness in mixture interaction assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion.
Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.
Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.
Treating mixture interaction assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.
Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.
Result / experiment / study / body of evidence, as applicable
Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies
OECD; OHAT; FDA Biomarker; EMA E16
https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics
Candidate
Pending owner approval; CC BY 4.0 recommended for OBO compatibility.
Source content preserved verbatim; computation metadata is a generated default and must be curated before normative use.
Established within specialty
Established
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 56
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Endpoint Validity; Endpoint Reliability; Endpoint Responsiveness", "Common Misinterpretations": "Treating outcome relevance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Outcome Relevance", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of outcome relevance.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses outcome relevance using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in outcome relevance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
56
1adca583ee04c245cdae3428bd050f3ed8796d29bf4b214b5880769da8977172
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 140
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Incremental Diagnostic Value; Prognostic Discrimination; Prognostic Calibration", "Common Misinterpretations": "Treating reclassification improvement as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Reclassification Improvement", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of reclassification improvement.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses reclassification improvement using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reclassification improvement can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
140
74d83029c18dff3247c8fb9baae1d6aba52c9b7328d218dd0210458cbd5ad7d5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 428
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Prespecified Analysis Adherence; Intervention Description Completeness; Comparator Description Completeness", "Common Misinterpretations": "Treating outcome definition completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Outcome Definition Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of outcome definition are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses outcome definition completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in outcome definition completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
428
129e8fff9350e8d3fb88d0a300e3816f498afcc03cf17b3f6037da094441e4fe
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 209
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Setting Applicability; Intervention Applicability; Comparator Applicability", "Common Misinterpretations": "Treating population applicability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Population Applicability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of population applicability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses population applicability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in population applicability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
209
9073b7f8b73fc788aa084d5c4ac174d69482c2a1439b22f86a39e465503328bc
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 462
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Interaction Assessment Adequacy; Measurement Error Correction; Calibration of Statistical Predictions", "Common Misinterpretations": "Treating overadjustment bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Overadjustment Bias Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The probability or degree that overadjustment bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses overadjustment bias risk using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in overadjustment bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
462
20f394b3ef37b7c0b5d8426ebcce5fe9f9ef841bbb2622101174fbb411ab7957
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 427
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Negative-Result Reporting; Deviations-from-Protocol Transparency; Reproducibility Information Completeness", "Common Misinterpretations": "Treating null-result interpretability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Null-Result Interpretability", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of null-result interpretability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses null-result interpretability using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in null-result interpretability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
427
0e7be34e4db71c86201d878decffa3c8da8ca62a1c8481dc3779d3c4caa454b1
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 294
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Limit of Quantification; Analytical Measurement Range; Reportable Range", "Common Misinterpretations": "Treating linearity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Linearity", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of linearity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses linearity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in linearity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
294
511e3fb111714ad52a916859d22dfd9a9aa36539d615e5b295fdf52500ac32e5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 116
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Residual Confounding Risk; Exchangeability Plausibility; Positivity Adequacy", "Common Misinterpretations": "Treating unmeasured confounding sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Unmeasured Confounding Sensitivity", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of unmeasured confounding sensitivity.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses unmeasured confounding sensitivity using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in unmeasured confounding sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
116
50677da95ed69c14f1f1ec5786cb5c4d590d6e339c1a13fdedfd18d811ffc8f4
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 274
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Analytical Validity; Trueness; Analytical Precision", "Common Misinterpretations": "Treating accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Accuracy", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The closeness of accuracy to the accepted reference or true value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses accuracy using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
274
35d988975cf78a06cb70788f89a4196551552de36a3cf47e57d2d96fc8cf41fd
editor note
Ontology-editor note.
BEMO:3200018
Context Specific Protocol Required
Controlled BEMO FormulaStatus value: ContextSpecificProtocolRequired.
BEMO:4000018
ContextSpecificProtocolRequired
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 208
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Comparator Applicability; Spectrum Representativeness; Subgroup Consistency", "Common Misinterpretations": "Treating outcome applicability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Outcome Applicability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of outcome applicability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses outcome applicability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in outcome applicability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
208
5c0db0a0257f7b171655d97624fdf4764d1fabad77716300abc92c2debe71f7d
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 12
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Rescue Experiment Support; Target Engagement Evidence; On-Target Specificity", "Common Misinterpretations": "Treating epistasis support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Epistasis Support", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The magnitude and credibility of independent evidence supporting epistasis.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses epistasis support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in epistasis support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
12
cd57f610728234561644257a8387987f0aa8934688593c5fbd817b3779032054
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 190
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Experimental Batch Randomization; Negative-Control Performance; Vehicle-Control Validity", "Common Misinterpretations": "Treating positive-control performance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Positive-Control Performance", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of positive-control performance.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses positive-control performance using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in positive-control performance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
190
98f43a472c5bc9ced611ad7ff89ac485c18a5c20ba0a8428164700738b649f75
Quantitative Ratio Computation
Controlled BEMO ComputationMode value: QuantitativeRatioComputation.
BEMO:4000009
QuantitativeRatioComputation
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 27
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Molecular-Phenotypic Concordance; Loss-of-Function Validation; Gain-of-Function Validation", "Common Misinterpretations": "Treating perturbational validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Perturbational Validation", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of perturbational validation.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses perturbational validation using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in perturbational validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
27
e56c259480c9a3f5c7efe467a62a9721824469292b9aeb900d1804f2bbfb7983
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 316
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Steady-State Validity; Perturbation Prediction Accuracy; Emergent-Property Reproducibility", "Common Misinterpretations": "Treating flux-balance consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Flux-Balance Consistency", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The degree of agreement in flux-balance across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses flux-balance consistency using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in flux-balance consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
316
f8d3b19efd68e0150e437c71bac90ab2bb80f33877b4a0e74bc52542a0fd3c26
2026-08-02
BEMO Project
A provisional, computable ontology of pure biomedical evidence research metrics. Each source metric is represented as an OWL class and linked to a versioned computation specification.
2026-08-02
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
Biomedical Evidence Metrics Ontology (BEMO)
DRAFT. This namespace is provisional, the ontology has not been reviewed or accepted by the OBO Foundry, and the final open license requires owner approval.
0.1.0
Pending owner approval; CC BY 4.0 recommended for OBO Foundry compatibility.
Computation profile for Care-Pathway Independence
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 181
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Attrition Accounting in Animal Studies; Experimental Batch Randomization; Positive-Control Performance", "Common Misinterpretations": "Treating exclusion-criteria prespecification as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Exclusion-Criteria Prespecification", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of exclusion-criteria prespecification.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses exclusion-criteria prespecification using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in exclusion-criteria prespecification can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
181
9718ea523adf00e3cd776b8e22423abc28cf080bd659786c7568ddd524ab9208
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 180
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Housing and Husbandry Control; Intervention Fidelity in Animal Studies; Humane Endpoint Appropriateness", "Common Misinterpretations": "Treating environmental standardization as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Environmental Standardization", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of environmental standardization.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses environmental standardization using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in environmental standardization can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
180
432e96e9d5c02fcda1a353b0ddb2e18d1019ff594f33c248eedfbe337bf31fae
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 255
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Call-Rate Completeness; Batch-Effect Control; Population Stratification Control", "Common Misinterpretations": "Treating hardy–weinberg equilibrium compatibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Hardy–Weinberg Equilibrium Compatibility", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of hardy–weinberg equilibrium compatibility.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses hardy–weinberg equilibrium compatibility using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in hardy–weinberg equilibrium compatibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
255
158d6abc0d325612b86f33eeb0aa8e0dbdfd4460dc9c1934f90601f28ed25284
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 401
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Direct Replication Success; Conceptual Replication Success", "Common Misinterpretations": "Treating independent replication strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Independent Replication Strength", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The magnitude and credibility of independent evidence supporting independent replication.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses independent replication strength using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in independent replication strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
401
3837461a84fca58515b90e02475b1cf195c845c2cf4704aa9d0c7c2b412d9406
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 329
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Module Stability; Pathway Topology Support; Knowledge-Graph Evidence Completeness", "Common Misinterpretations": "Treating pathway enrichment consistency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Pathway Enrichment Consistency", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The degree of agreement in pathway enrichment across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses pathway enrichment consistency using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in pathway enrichment consistency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
329
d7e323b8eddcaff29fe662695a2fa9bae2a0f04120cfdf422fa3e2d1cd3b117e
required inputs text
Semicolon-delimited required input roles.
BEMO:3100006
has required input specification
Connects a computation specification to a required input specification.
BEMO:3000004
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 155
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Consistency; Evidence Precision; Evidence Coherence", "Common Misinterpretations": "Treating evidence directness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Directness", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence directness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses evidence directness using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence directness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
155
880de5f192c9dc60c14576e428f1c46fbdcda3180afe950bcde297920e65c6df
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 24
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "On-Target Specificity; Biological Gradient; Homeostatic Compensation Assessment", "Common Misinterpretations": "Treating off-target liability evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Off-Target Liability Evidence", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of off-target liability evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses off-target liability evidence using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in off-target liability evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
24
688a2bdbd29ad4b59eca6350cf3462f2bc158ee3d70ef4de59f338b6e6c797a6
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 89
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Pathology Confirmation; Cellularity Adequacy; Necrosis Burden", "Common Misinterpretations": "Treating tumor purity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Tumor Purity", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of tumor purity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses tumor purity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in tumor purity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
89
e7836bb1f6728d199c8634dbdbf560dae6243a3932920cc82164935a37c9180e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 358
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Carcinogenicity Evidence Strength; Developmental Toxicity Evidence Strength; Immunotoxicity Evidence Strength", "Common Misinterpretations": "Treating reproductive toxicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Reproductive Toxicity Evidence Strength", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The magnitude and credibility of independent evidence supporting reproductive toxicity evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses reproductive toxicity evidence strength using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reproductive toxicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
358
1725f0845040d88689ff2869cde10ea44454b5cfaa1c303d02071f7ff5ccb751
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 363
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Dose Proportionality; Metabolite Coverage; Drug–Drug Interaction Evidence", "Common Misinterpretations": "Treating time–concentration profile adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Time–Concentration Profile Adequacy", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The extent to which time–concentration profile is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Method- and analyte-specific physical units", "What It Measures": "Assesses time–concentration profile adequacy using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in time–concentration profile adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
363
8bfcea065494f7118cb14f95a9459e3f35fa404437872877955c4fdff069147a
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 280
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Site-to-Site Assay Portability; Preanalytical Robustness; Postanalytical Integrity", "Common Misinterpretations": "Treating batch-effect sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Batch-Effect Sensitivity", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of batch-effect sensitivity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses batch-effect sensitivity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in batch-effect sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
280
7222c30713cc035f86bae406549aaf0d4751a919d50ae68b917efd9629244d28
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 165
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Counterevidence Strength; Publication Bias Risk; Selective Nonreporting Risk", "Common Misinterpretations": "Treating missing evidence risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Missing Evidence Risk", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The probability or degree that missing evidence introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses missing evidence risk using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in missing evidence risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
165
ffec67315125f9639a3601d51e352906c34b1c9847a274942f934888c1188977
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 20
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Mechanistic Coherence; Mechanistic Causality Strength; Pathway-Level Support", "Common Misinterpretations": "Treating mechanistic specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Mechanistic Specificity", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of mechanistic specificity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses mechanistic specificity using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in mechanistic specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
20
29074cf2dac765628d0a639b922f2e3313846aa414af9c4ef760ea93c86819ca
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 59
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Predictive Biomarker Validity; Diagnostic Biomarker Validity; Pharmacodynamic Biomarker Validity", "Common Misinterpretations": "Treating prognostic biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Prognostic Biomarker Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which prognostic biomarker supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses prognostic biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in prognostic biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
59
2dd804ff14e01e089d83a5d2d6ef790d24cf90ca54d3426049116e4e7aeb095f
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 152
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Coherence; Evidence Sufficiency; Evidence Completeness", "Common Misinterpretations": "Treating evidence consensus strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Consensus Strength", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The magnitude and credibility of independent evidence supporting evidence consensus.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses evidence consensus strength using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence consensus strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
152
f207ed4da0508b76cb7e86f43fc62efc4861d327760c860608fc73ff3f4a6da8
Evidence Certainty and Synthesis metric
Category of biomedical evidence metrics concerned with evidence certainty and synthesis.
bemo
BEMO:1100006
Candidate
Operational Definition Required
Controlled BEMO ComputationReadinessStatus value: OperationalDefinitionRequired.
BEMO:4000014
OperationalDefinitionRequired
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 73
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Time-to-Fixation Adequacy; Preservation Adequacy; Storage Temperature Control", "Common Misinterpretations": "Treating fixation adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Fixation Adequacy", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "The extent to which fixation is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses fixation adequacy using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in fixation adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
73
c08e61688a36fec680e9073fc8888dbfd98852919e3f7ccf7c3180ed49d1e921
depends on metric
Represents a curated computational dependency between metrics.
BEMO:3000016
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 156
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Coverage; Evidence Robustness; Evidence Triangulation Strength", "Common Misinterpretations": "Treating evidence freshness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Freshness", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence freshness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses evidence freshness using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence freshness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
156
2dc879dd1ad8ed3ebf228b83983e7de95e06b9c72131b584ccaa785f551eea2b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 210
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "External Validity; Sampling Frame Adequacy; Transportability", "Common Misinterpretations": "Treating population representativeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Population Representativeness", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of population representativeness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses population representativeness using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in population representativeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
210
efa730876069251ed2dde336a59f381b8c109c4ee429da0b68bf8bae61b0ba48
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 50
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Endpoint Reliability; Minimal Clinically Important Difference Validity", "Common Misinterpretations": "Treating endpoint responsiveness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Endpoint Responsiveness", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of endpoint responsiveness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses endpoint responsiveness using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in endpoint responsiveness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
50
3cb1f45a15a5daa5f6b98c2d745a57440f42e5840e757c82a744e9f010f76df0
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 350
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Human-Relevance of Toxicological Evidence", "Common Misinterpretations": "Treating mixture interaction assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Mixture Interaction Assessment", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of mixture interaction assessment.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses mixture interaction assessment using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in mixture interaction assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
350
241ffa4364ee9325fbb3832728344d8484375a1aca172f524b6f871b9dc07687
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 192
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Technical Replicate Adequacy; Randomization in Experimental Allocation; Blinding in Experimental Assessment", "Common Misinterpretations": "Treating sample size justification as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Sample Size Justification", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of sample size justification.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses sample size justification using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in sample size justification can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
192
9c4289123cb7c0f9d87fa9453151e360a9b6781cd9fbb6f6b26a09f4543b0dab
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 263
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Contamination Burden; Sex Concordance; Variant Call Quality", "Common Misinterpretations": "Treating sample identity concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Sample Identity Concordance", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "The degree of agreement in sample identity across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses sample identity concordance using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in sample identity concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
263
cf22f061ff0a6f0403c89f061c9c86c0fdc0e5186e885958801878829549f2cd
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 303
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Calibration Traceability; Cutoff Validity; Measurement Uncertainty", "Common Misinterpretations": "Treating reference interval validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Reference Interval Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The degree to which reference interval supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses reference interval validity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reference interval validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
303
3706e41c2a6c519ccdfb2643facf47ec1d2a45cc4b5b902e54040a5a53624b7e
Pharmacology and Toxicology metric
Category of biomedical evidence metrics concerned with pharmacology and toxicology.
bemo
BEMO:1100013
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 79
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Protein Integrity; Chain-of-Custody Integrity; Matched-Sample Integrity", "Common Misinterpretations": "Treating microbial contamination as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Microbial Contamination", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of microbial contamination.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses microbial contamination using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in microbial contamination can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
79
a7e7bf9a47cefd808a5ba122413c735b8b095ffc84592451f9857f4fe080e389
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 452
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Nonlinearity Assessment; Overadjustment Bias Risk; Measurement Error Correction", "Common Misinterpretations": "Treating interaction assessment adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Interaction Assessment Adequacy", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The extent to which interaction assessment is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses interaction assessment adequacy using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in interaction assessment adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
452
7942499ba472159741e9410722d92887d17131b3d661f12b98b8ec8f2d490791
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 488
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Blinding Integrity; Detection Bias Risk; Attrition Bias Risk", "Common Misinterpretations": "Treating performance bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Performance Bias Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that performance bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses performance bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in performance bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
488
67d1ca80d7087e3cfba7c1c9f64d45d3f4094887ac27aef09ef61bfb6d49d732
Computation profile for Prognostic Transportability
0.1.0
Apply a versioned metric-specific computation protocol consistent with the source definition, criteria, and methods.
BEMO-Expression-JSON
{"language":"BEMO-Expression-JSON","operator":"external_protocol","protocolRef":"REQUIRED"}
evidence_records; assessment_context; operational_definition; computation_protocol_version
weights; thresholds; reference_standard; expert_adjudication
xsd:anySimpleType
Metric-specific continuous, categorical, or ordinal scale
Not specified in source; must be defined and versioned before composite use.
None by default; any normalization must be justified, versioned, and validated.
Must be declared before computation; report missingness; no silent imputation; perform sensitivity analysis when material.
true
Confidence interval, credible interval, bootstrap distribution, inter-rater reliability, or sensitivity analysis as scientifically applicable.
true
Prespecify, justify, version, and sensitivity-test thresholds; do not derive and evaluate on the same data without correction.
Source provenance; input validation; duplicate control; uncertainty reporting; independent or orthogonal validation where applicable.
As applicable; mandatory for probabilistic or normalized outputs.
Source maturity: Established; BEMO computation profile requires independent validation.
Draft
Controlled BEMO ApprovalStatus value: Draft.
BEMO:4000027
Draft
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 468
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Type I Error Control; Multiplicity Control; Model Specification Adequacy", "Common Misinterpretations": "Treating type ii error risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Type II Error Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The probability or degree that type ii error introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses type ii error risk using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in type ii error risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
468
de4287424d92a6f5a119531d8abefacf3d79beff40059540425ff1afc5f1021d
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 307
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Commutability; Reagent Lot Consistency; Instrument Drift", "Common Misinterpretations": "Treating sample stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Sample Stability", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of sample stability.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses sample stability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in sample stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
307
42f10127b0bb2f21e12bfaaba11ee2a6c8335cd39935669d6733741eadfc7015
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 309
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Accuracy; Analytical Precision; Repeatability", "Common Misinterpretations": "Treating trueness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Trueness", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of trueness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses trueness using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in trueness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
309
9094cf98fb975c34384185560acacc946dab62ebd8f1815100dcaac099d9df1b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 149
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Precision; Evidence Consensus Strength; Evidence Sufficiency", "Common Misinterpretations": "Treating evidence coherence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Coherence", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence coherence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses evidence coherence using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence coherence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
149
573628981de9da600edff4f13001018973ec4403d9ad58179a02f6958d0404dd
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 434
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Processed-Data Availability; Negative-Result Reporting; Null-Result Interpretability", "Common Misinterpretations": "Treating quality-control reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Quality-Control Reporting Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of quality-control reporting are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses quality-control reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in quality-control reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
434
b53b586f994fd48c27b9acc1c6fff59b966085610e593ec05c03735d90e329b2
validation status text
Human-readable validation status.
BEMO:3100030
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 344
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Pharmacological Target Validity; Pharmacokinetic Adequacy; Pharmacodynamic Adequacy", "Common Misinterpretations": "Treating exposure–response relationship as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Exposure–Response Relationship", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of exposure–response relationship.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses exposure–response relationship using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in exposure–response relationship can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
344
882d0346f9d5903a7e9b6a2b5ae992e9e24bd2be9c7e249168f2f9e157dc0354
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 120
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Test-Timing Appropriateness; Incremental Diagnostic Value; Reclassification Improvement", "Common Misinterpretations": "Treating comparative test accuracy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Comparative Test Accuracy", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "The closeness of comparative test to the accepted reference or true value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses comparative test accuracy using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in comparative test accuracy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
120
01f4b65430acfaca391e550bb1550450d424a1dd47aebe4ed0f6dd9d60fe28ab
Measurement and Assay Analytical Validity metric
Category of biomedical evidence metrics concerned with measurement and assay analytical validity.
bemo
BEMO:1100011
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 77
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Hemolysis Burden; Icterus Interference; RNA Integrity", "Common Misinterpretations": "Treating lipemia burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Lipemia Burden", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of lipemia burden.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses lipemia burden using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in lipemia burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
77
572ca1cf50802a5447be4c1286cdad2930a7ec08e97e7b9af663a3cb66e1d788
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 302
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Dynamic Range; Dilution Integrity; Matrix Effect", "Common Misinterpretations": "Treating recovery as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Recovery", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of recovery.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses recovery using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in recovery can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
302
5e396bd1803ce209b22910e5705b5bf0dc26bd235bba5ac5c904b2514e3f7ea8
source sheet
Name of the source worksheet.
BEMO:3200015
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 276
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Trueness; Repeatability; Intermediate Precision", "Common Misinterpretations": "Treating analytical precision as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Analytical Precision", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The closeness of repeated estimates or measurements and the narrowness of uncertainty around analytical.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses analytical precision using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in analytical precision can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
276
7c4987f5f8c3c9cca41bb28896f2392e5c664fc505d3f5db2c63ad7c10e1d04d
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 130
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Positive Predictive Value; Positive Likelihood Ratio; Negative Likelihood Ratio", "Common Misinterpretations": "Treating negative predictive value as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Negative Predictive Value", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of negative predictive value.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses negative predictive value using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in negative predictive value can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
130
9f42e5362ad2e55445a028ec6d3cc9c062886d8c9272f983b0c1acec345502f7
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 435
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Metadata Completeness; Processed-Data Availability; Quality-Control Reporting Completeness", "Common Misinterpretations": "Treating raw-data availability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Raw-Data Availability", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of raw-data availability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses raw-data availability using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in raw-data availability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
435
2fcf4f96e7ca202f61d3de161ead38ea506d30ac8f05b66ac8a84b27f3da645e
confidence interval required
Whether interval estimation is required.
BEMO:3100018
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 251
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Relatedness Control; Normalization Adequacy; Transcript Quantification Reliability", "Common Misinterpretations": "Treating differential expression robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Differential Expression Robustness", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of differential expression robustness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses differential expression robustness using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in differential expression robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
251
e70802b2991914211c0dcd91213c02319dd673fc93963b9e2064b8bca944c528
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 482
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Follow-up Completeness", "Common Misinterpretations": "Treating differential follow-up risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Differential Follow-up Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that differential follow-up introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses differential follow-up risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in differential follow-up risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
482
ccfe7e445aeb97f5f6fbabfec00706b7cd315705f68e05e908e47860b40d8ad4
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 295
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Dilution Integrity; Interference Susceptibility; Cross-Reactivity", "Common Misinterpretations": "Treating matrix effect as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Matrix Effect", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of matrix effect.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses matrix effect using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in matrix effect can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
295
e901d4ca31939164c2ecbe0d0bed735b156c8e09a3c3f28444f5359947798b49
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 442
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Estimate Precision; Statistical Power; Type I Error Control", "Common Misinterpretations": "Treating confidence interval compatibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Confidence Interval Compatibility", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The justified degree of certainty assigned to confidence interval compatibility given the quantity, quality, consistency, and limitations of supporting evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses confidence interval compatibility using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in confidence interval compatibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
442
4b44e65362db0ef8f1ef0ae6eb87b631fd22642d5790f792ab6c3d6a9a693508
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 10
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Pathway-Level Support; Tissue Specificity; Spatial Biological Concordance", "Common Misinterpretations": "Treating cell-type specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Cell-Type Specificity", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cell-type specificity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses cell-type specificity using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cell-type specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
10
8fcb10135485d0dff26bef51519f2178a09a543c766f354a728ca79b2c39c582
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 256
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Duplicate Read Burden; Contamination Burden; Sample Identity Concordance", "Common Misinterpretations": "Treating library complexity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Library Complexity", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of library complexity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses library complexity using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in library complexity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
256
fb650cc4b9fc46d550895b59d1656dada1a7ec625426686e0a708ba75b6fc5f0
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 470
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Early Stopping Bias Risk; Exposure Classification Validity; Intervention Classification Validity", "Common Misinterpretations": "Treating adherence integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Adherence Integrity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of adherence integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses adherence integrity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in adherence integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
470
9a20a1494ccbc46d7bc0acd4d260241ebc29edd15b77d8ccc525c9f0dba8e8de
lifecycle status
Term lifecycle status.
BEMO:3200016
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 16
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Mechanistic Specificity; Pathway-Level Support; Cell-Type Specificity", "Common Misinterpretations": "Treating mechanistic causality strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Mechanistic Causality Strength", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The magnitude and credibility of independent evidence supporting mechanistic causality.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses mechanistic causality strength using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in mechanistic causality strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
16
114851a566f6667271761fd50b701ab810f623c4ee0ecdabe88d5bbacc5941db
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 131
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Calibration-in-the-Large; Prognostic Added Value; Prognostic Transportability", "Common Misinterpretations": "Treating observed-to-expected ratio as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Observed-to-Expected Ratio", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of observed-to-expected ratio.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses observed-to-expected ratio using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in observed-to-expected ratio can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
131
cdf33871a90f7426da799fbf8148672bff9e925c7981ff58372c39511bd4050e
Template Computable
Controlled BEMO ComputationReadinessStatus value: TemplateComputable.
BEMO:4000017
TemplateComputable
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 236
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Computational Variant Evidence; Variant Phase Evidence; Loss-of-Function Mechanism Validity", "Common Misinterpretations": "Treating phenotypic specificity for variant as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Phenotypic Specificity for Variant", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of phenotypic specificity for variant.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses phenotypic specificity for variant using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in phenotypic specificity for variant can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
236
ae0b81578c2487ac58cc25c0925e5e5b5580990b17b1391fc512e266a26a2788
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 371
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Isotope Pattern Fidelity; Post-Translational Modification Localization Confidence; Proteoform Identification Confidence", "Common Misinterpretations": "Treating fragmentation spectrum quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Fragmentation Spectrum Quality", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of fragmentation spectrum quality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses fragmentation spectrum quality using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in fragmentation spectrum quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
371
4295f539b87882bc52461f7bd190419d6afdf7864927b9a24efe1e5f65c9fcd7
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 385
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Peptide-Spectrum Match Quality; Proteome Coverage; Sequence Coverage", "Common Misinterpretations": "Treating protein inference reliability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Protein Inference Reliability", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of protein inference reliability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses protein inference reliability using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in protein inference reliability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
385
601ad0978c475da6065d2502b00169f5521ae97cb0543ad15ad88f49903cb27e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 421
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Participant Flow Completeness; Statistical Methods Reporting Completeness; Missing-Data Reporting Completeness", "Common Misinterpretations": "Treating harms reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Harms Reporting Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of harms reporting are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses harms reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in harms reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
421
15ffc51fd66113f5e5c480a7127c592104dbff4a2e6faafcfafb509cd725ea07
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 356
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Selectivity Profile; Efficacy Reproducibility; Therapeutic Window Evidence", "Common Misinterpretations": "Treating potency reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Potency Reproducibility", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The degree to which potency reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses potency reproducibility using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in potency reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
356
2199b92ac6c23ce1866fa6fdae782c65a478ec51834e2f0193f79d608ea4c3e7
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 76
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Lipemia Burden; RNA Integrity; DNA Integrity", "Common Misinterpretations": "Treating icterus interference as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Icterus Interference", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of icterus interference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses icterus interference using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in icterus interference can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
76
d0f3346c9caa8abb9b49bded0ab6e26e5f13d5a8fb5a20971f64067dd82054ea
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 33
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Spatial Biological Concordance; Cross-Species Biological Concordance; Phenotypic Concordance", "Common Misinterpretations": "Treating temporal biological concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Temporal Biological Concordance", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The degree of agreement in temporal biological across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses temporal biological concordance using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in temporal biological concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
33
57d68c3c49543e85a4a6ee1a26a1705e9e0e05df169673d5a8f0a00632bfa39b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 85
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Icterus Interference; DNA Integrity; Protein Integrity", "Common Misinterpretations": "Treating rna integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "RNA Integrity", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of rna integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses rna integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in rna integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
85
98ac509f170b3163d862e8e08cc16dad8b915df5508f4e7394a8627d881f7350
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 123
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Diagnostic Specificity; Positive Predictive Value", "Common Misinterpretations": "Treating diagnostic sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Diagnostic Sensitivity", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of diagnostic sensitivity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses diagnostic sensitivity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in diagnostic sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
123
4ea75714eea731a1e4644a00db0476f8628bf028934f92bd9e4029ee31ddf173
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 96
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Positivity Adequacy; No-Interference Plausibility; Correct Temporal Ordering", "Common Misinterpretations": "Treating consistency assumption plausibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Consistency Assumption Plausibility", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The degree of agreement in consistency assumption plausibility across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses consistency assumption plausibility using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in consistency assumption plausibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
96
1bec3d5308662b9b5fad28978574efe7573d32766cf45ac34286863d2d2267a8
measurement criteria
Preserves the source measurement criteria.
BEMO:3200004
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 117
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Diagnostic Odds Ratio; Partial Area Under the Receiver Operating Characteristic Curve; Threshold Validity", "Common Misinterpretations": "Treating area under the receiver operating characteristic curve as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Area Under the Receiver Operating Characteristic Curve", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of area under the receiver operating characteristic curve.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses area under the receiver operating characteristic curve using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in area under the receiver operating characteristic curve can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
117
5c4feaae4c2a565042108ac4cb807678f9be09f3e84e62f596afd3338e3ceec6
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 372
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Spectral Library Match Quality; Extraction Recovery; Derivatization Efficiency", "Common Misinterpretations": "Treating internal standard performance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Internal Standard Performance", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of internal standard performance.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses internal standard performance using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in internal standard performance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
372
904dce77d19fcc637d6eae56d2d92b19b7052fd578de120e2a18bb08d00f2f54
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 128
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Reference Standard Validity; Verification Bias Risk; Differential Verification Bias Risk", "Common Misinterpretations": "Treating index-test blinding as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Index-Test Blinding", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of index-test blinding.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses index-test blinding using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in index-test blinding can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
128
3bc623a30f8cb3956f6d71196888c1021137e72c87604513c9fc1e352f98bb4c
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 147
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Triangulation Strength; Missing Evidence Risk; Publication Bias Risk", "Common Misinterpretations": "Treating counterevidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Counterevidence Strength", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The magnitude and credibility of independent evidence supporting counterevidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses counterevidence strength using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in counterevidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
147
d4c1d3bbceb7e86e058c678a7bbc386dd5980f93dcbc9a60ebc6c6d39231fcdb
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 87
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Cold Ischemia Control; Fixation Adequacy; Preservation Adequacy", "Common Misinterpretations": "Treating time-to-fixation adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Time-to-Fixation Adequacy", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "The extent to which time-to-fixation is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses time-to-fixation adequacy using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in time-to-fixation adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
87
49559a7ae628351af963a27a55893cf501098be34876ce35bdc2d4969fd4f94b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 286
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Recovery; Matrix Effect; Interference Susceptibility", "Common Misinterpretations": "Treating dilution integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Dilution Integrity", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of dilution integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses dilution integrity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in dilution integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
286
a31ef0b724223ebc3d1ad0dbe731b0b7823f6b347840ea43769f01a33866757b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 297
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Measurement Uncertainty; Commutability; Sample Stability", "Common Misinterpretations": "Treating method comparison agreement as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Method Comparison Agreement", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The degree of agreement in method comparison across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses method comparison agreement using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in method comparison agreement can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
297
8f2dfbddd616818e03a89990f48f31b1db69460f1919bb65cc71ee0d68b7b4f6
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 298
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Instrument Drift; Site-to-Site Assay Portability; Batch-Effect Sensitivity", "Common Misinterpretations": "Treating operator variability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Operator Variability", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of operator variability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses operator variability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in operator variability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
298
ae9e50b8f8857543ecb1eb315c23fc88dd3a028cfd40ee4377e1039431544ca0
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 142
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Patient Flow Integrity; Comparative Test Accuracy; Incremental Diagnostic Value", "Common Misinterpretations": "Treating test-timing appropriateness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Test-Timing Appropriateness", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "The extent to which test-timing is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses test-timing appropriateness using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in test-timing appropriateness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
142
ffb2707e3bed89aca5f5c790dc5a9b0312e89c5a1dba27ebac908e4597a355bd
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 145
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Index-Test Blinding; Differential Verification Bias Risk; Incorporation Bias Risk", "Common Misinterpretations": "Treating verification bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Verification Bias Risk", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "The probability or degree that verification bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses verification bias risk using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in verification bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
145
2f165758b70b1e482bfb8517e1a69a7d4e05ba84cd41dc36c2f346f991371034
has uncertainty specification
Connects a computation specification or assessment to an uncertainty specification.
BEMO:3000012
optional inputs text
Semicolon-delimited optional input roles.
BEMO:3100007
Measurement, Assay, and Biospecimen Quality metric
Metrics assessing analytical measurement performance, assay validity, biospecimen handling, and preanalytical quality.
bemo
BEMO:1000003
Candidate
minimum value
Minimum allowed or expected numeric value.
BEMO:3100010
metric value lexical form
Lexical representation of a metric assessment value.
BEMO:3100025
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 243
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Phenotypic Specificity for Variant; Loss-of-Function Mechanism Validity; Hotspot/Functional-Domain Evidence", "Common Misinterpretations": "Treating variant phase evidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Variant Phase Evidence", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of variant phase evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses variant phase evidence using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in variant phase evidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
243
e63350cbc225389693f86f1ceb6c24e47930b1a6b522574479bcaf1390376503
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 336
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Toxicological Mode-of-Action Support; Organ-Specific Toxicity Evidence; Genotoxicity Evidence Strength", "Common Misinterpretations": "Treating adverse outcome pathway support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Adverse Outcome Pathway Support", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The magnitude and credibility of independent evidence supporting adverse outcome pathway.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses adverse outcome pathway support using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in adverse outcome pathway support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
336
8935703e7b30761e4213b86ddafc1aecae0696f4523a201860d9e64493990a74
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 43
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Construct Validity; Predictive Biomarker Validity; Prognostic Biomarker Validity", "Common Misinterpretations": "Treating content validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Content Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which content supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses content validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in content validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
43
1ceb580fd3a672a7c14d44163bda2886dd03d14b99ce004f1db40aa485bb3e1b
assesses evidence object
Connects a metric assessment to the evidence object assessed.
BEMO:3000003
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 487
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Protocol Fidelity; Follow-up Completeness; Differential Follow-up Risk", "Common Misinterpretations": "Treating outcome ascertainment validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Outcome Ascertainment Validity", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The degree to which outcome ascertainment supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses outcome ascertainment validity using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in outcome ascertainment validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
487
654f73aad6650cd58d63f5a49ae9e034bfcb22952cbf7a018f652c070e1ba6c1
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 416
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Funding-Source Transparency; Data-Sharing Transparency; Code-Sharing Transparency", "Common Misinterpretations": "Treating conflict-of-interest transparency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Conflict-of-Interest Transparency", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of conflict-of-interest transparency.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses conflict-of-interest transparency using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in conflict-of-interest transparency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
416
e951e1891736a22e6ef317b85f2f1f3032c69473b10d7ac72b1ed8342656df25
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 402
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Result Reproducibility; Reanalysis Concordance; Data Provenance Completeness", "Common Misinterpretations": "Treating inferential reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Inferential Reproducibility", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The degree to which inferential reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses inferential reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in inferential reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
402
b49eb32e7bf118727039f0240b17aee8970c665e4a359a2bcecc87341188ecf1
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 399
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Independent Replication Strength; Conceptual Replication Success; Analytical Reproducibility", "Common Misinterpretations": "Treating direct replication success as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Direct Replication Success", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of direct replication success.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses direct replication success using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in direct replication success can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
399
e6b0ec1b803bb2c7cb8becab175dbb88ab91c9d8ccd3f713a814c836846ea547
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 382
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "False Discovery Rate Control; Protein Inference Reliability; Proteome Coverage", "Common Misinterpretations": "Treating peptide-spectrum match quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Peptide-Spectrum Match Quality", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of peptide-spectrum match quality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses peptide-spectrum match quality using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in peptide-spectrum match quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
382
69f7898b8265a7daaa45bd90d4d5387c92da6521721d3c4901a5c208a63f0169
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 441
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Decision-Curve Net Benefit; Fragility Index; Bayes Factor Evidence", "Common Misinterpretations": "Treating clinical relevance of effect as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Clinical Relevance of Effect", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of clinical relevance of effect.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses clinical relevance of effect using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in clinical relevance of effect can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
441
7e04bf59de003d9389f3d5e5a39c09dbec7b25b54fdd665e24d339ccf148fb17
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 111
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Confounding Risk; Unmeasured Confounding Sensitivity; Exchangeability Plausibility", "Common Misinterpretations": "Treating residual confounding risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Residual Confounding Risk", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The probability or degree that residual confounding introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Directed acyclic graphs; design emulation; balance diagnostics; negative controls; quantitative bias analysis; sensitivity and falsification analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses residual confounding risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in residual confounding risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
111
fef3c7d8a6ff0d0d4088a08abb2cfe7420127e6170b4998b7b5ce06d44124e81
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 343
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Potency Reproducibility; Therapeutic Window Evidence; Safety Margin Evidence", "Common Misinterpretations": "Treating efficacy reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Efficacy Reproducibility", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The degree to which efficacy reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses efficacy reproducibility using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in efficacy reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
343
e3d46320b779b7100a656dae6dcb02f7b9eae6762c1680f48b9313843cb64c84
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 334
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Dynamical Stability; Flux-Balance Consistency; Perturbation Prediction Accuracy", "Common Misinterpretations": "Treating steady-state validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Steady-State Validity", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The degree to which steady-state supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses steady-state validity using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in steady-state validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
334
3a349724361b2712b35c94baaaa135ce8e313533d3edafc463242f1cddef411b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 62
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Response Biomarker Validity; Individual-Level Surrogacy; Trial-Level Surrogacy", "Common Misinterpretations": "Treating surrogate endpoint validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Surrogate Endpoint Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which surrogate endpoint supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses surrogate endpoint validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in surrogate endpoint validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
62
880b1c4a9fa2a25e2876679a2a7e759365591c8d0d0ae431e9a6e89a9d327bd9
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 8
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Off-Target Liability Evidence; Homeostatic Compensation Assessment; Network Context Support", "Common Misinterpretations": "Treating biological gradient as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Biological Gradient", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of biological gradient.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses biological gradient using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biological gradient can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
8
6471dcf670c9dd05424f4606389decf6b51812d07c0a0680842d77bda957b0f3
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 108
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Consistency Assumption Plausibility; Correct Temporal Ordering; Causal Contrast Clarity", "Common Misinterpretations": "Treating no-interference plausibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "No-Interference Plausibility", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of no-interference plausibility.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Replicate dilution series; blank and spiked samples; reference materials; method-comparison studies; predefined CLSI/ISO acceptance criteria.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses no-interference plausibility using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in no-interference plausibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
108
8aa9a1ba76fa5fd8217fff1b4352656d262602a1040dc8618b138b2e37a93d9a
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 193
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Species Appropriateness; Age Appropriateness; Genetic Background Control", "Common Misinterpretations": "Treating sex as a biological variable adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Sex as a Biological Variable Adequacy", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The extent to which sex as a biological variable is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses sex as a biological variable adequacy using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in sex as a biological variable adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
193
67d1bc84f2c2d0324ee734ed3d2df1c72cd4820b21290b5a8b136d75863a4866
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 40
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Clinical Validity; Biomarker Clinical Relevance; Context-of-Use Validity", "Common Misinterpretations": "Treating biomarker-outcome association strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Biomarker-Outcome Association Strength", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The magnitude and credibility of independent evidence supporting biomarker-outcome association.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses biomarker-outcome association strength using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biomarker-outcome association strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
40
d54499df144e3534026302c035690746b731c0abf25107f5fccdd8bada0bd79a
Research Transparency and Reporting Completeness metric
Category of biomedical evidence metrics concerned with research transparency and reporting completeness.
bemo
BEMO:1100016
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 467
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Statistical Power; Type II Error Risk; Multiplicity Control", "Common Misinterpretations": "Treating type i error control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Type I Error Control", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of type i error control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses type i error control using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in type i error control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
467
a3eba05a4613d95dc796e3df689756346c4f78ac2fda1dcb4e8cf4b813a22638
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 423
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Code-Sharing Transparency; Metadata Completeness; Raw-Data Availability", "Common Misinterpretations": "Treating materials-and-reagents reporting completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Materials-and-Reagents Reporting Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of materials-and-reagents reporting are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses materials-and-reagents reporting completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in materials-and-reagents reporting completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
423
fc72795f141f7cb81bc8cc86190a0c7cf2a5f2f6065d3b22f9bc88dd2e13080e
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 204
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Population Representativeness; Sampling Frame Adequacy", "Common Misinterpretations": "Treating external validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "External Validity", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "The degree to which external supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses external validity using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in external validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
204
291600cbcf719af0c4e82629e00cccabf1b55b19eff672c1fb6c5c7fe2ec4338
Diagnostic and Prognostic Evidence metric
Category of biomedical evidence metrics concerned with diagnostic and prognostic evidence.
bemo
BEMO:1100005
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 69
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Microbial Contamination; Matched-Sample Integrity", "Common Misinterpretations": "Treating chain-of-custody integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Chain-of-Custody Integrity", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of chain-of-custody integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses chain-of-custody integrity using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in chain-of-custody integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
69
38ecb99023f4b12b1bfca7d66221917ae384546ef170c868cd32fb2eb70f3f40
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 13
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Loss-of-Function Validation; Rescue Experiment Support; Epistasis Support", "Common Misinterpretations": "Treating gain-of-function validation as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Gain-of-Function Validation", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of gain-of-function validation.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses gain-of-function validation using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in gain-of-function validation can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
13
5c02cfac370a9d2c1514479546d6e291e60bffe1758412ceec02bbf8edaab2dc
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 308
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Operator Variability; Batch-Effect Sensitivity; Preanalytical Robustness", "Common Misinterpretations": "Treating site-to-site assay portability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Site-to-Site Assay Portability", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of site-to-site assay portability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses site-to-site assay portability using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in site-to-site assay portability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
308
c0e7317023f6c1fc290d021e4b38163503a72ee411cc671fe6aae25ee84d1ee5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 335
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Parameter Sensitivity; Practical Identifiability; Dynamical Stability", "Common Misinterpretations": "Treating structural identifiability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Structural Identifiability", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of structural identifiability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses structural identifiability using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in structural identifiability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
335
36fbdf592582d2c048b53d43d95edfca8afe5d6c7c93d352f32f2a7d7644cfbe
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 300
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Batch-Effect Sensitivity; Postanalytical Integrity", "Common Misinterpretations": "Treating preanalytical robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Preanalytical Robustness", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of preanalytical robustness.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses preanalytical robustness using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in preanalytical robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
300
b76350bac6098346689193598ba9a67fbdcf0cc16c9a9abb5e49d8238a97798b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 218
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "De Novo Evidence Strength; Case-Level Evidence Strength; Case-Control Evidence Strength", "Common Misinterpretations": "Treating allelic evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Allelic Evidence Strength", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The magnitude and credibility of independent evidence supporting allelic evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses allelic evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in allelic evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
218
2dd4a685750034c6175101f0128d889784cf621fed0cf93a74098e6b4879332b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 351
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Safety Margin Evidence; Lowest-Observed-Adverse-Effect Level Robustness; Benchmark Dose Reliability", "Common Misinterpretations": "Treating no-observed-adverse-effect level robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "No-Observed-Adverse-Effect Level Robustness", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of no-observed-adverse-effect level robustness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses no-observed-adverse-effect level robustness using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in no-observed-adverse-effect level robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
351
f5f71661574e809d9eb1b7625ffde603665b4c0d25ddbb968f6b8d1a88c38a42
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 493
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Attrition Bias Risk; Protocol Deviation Risk; Contamination Risk", "Common Misinterpretations": "Treating selective outcome reporting risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Selective Outcome Reporting Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that selective outcome reporting introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses selective outcome reporting risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in selective outcome reporting risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
493
bcdaa5f06df0a5ab7941c490a740787093ffc20f433c48125d9c2e174850f7fd
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 266
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Single-Cell Doublet Burden; Single-Cell Viability; Single-Cell Feature Detection Rate", "Common Misinterpretations": "Treating single-cell ambient rna burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Single-Cell Ambient RNA Burden", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of single-cell ambient rna burden.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses single-cell ambient rna burden using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in single-cell ambient rna burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
266
1685a23504d9a42b00cb2ea9bdca0c2daeec5f3be11cc0d479a9acb5349c658a
metric identifier
Stable OBO-style CURIE assigned to a metric.
BEMO:3100001
common misinterpretations
Preserves the source warnings about misinterpretation.
BEMO:3200006
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 457
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Multiplicity Control; Model Fit; Residual Diagnostics Adequacy", "Common Misinterpretations": "Treating model specification adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Model Specification Adequacy", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The extent to which model specification is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses model specification adequacy using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in model specification adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
457
c5d48404d8de034fc4aa45bb9f64638b1d861e61321423aa5f11d4f63cc6783b
metric computation process
A process that applies a metric computation specification to validated inputs and produces a metric assessment.
bemo
BEMO:0000300
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 124
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Diagnostic Sensitivity; Positive Predictive Value; Negative Predictive Value", "Common Misinterpretations": "Treating diagnostic specificity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Diagnostic Specificity", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of diagnostic specificity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses diagnostic specificity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in diagnostic specificity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
124
14be8f0b88f11902e6b1562a38b9cb710840858fc59f748048f96b74c8f6c2b2
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 445
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Residual Diagnostics Adequacy; Variance Estimation Validity; Missing-Data Mechanism Plausibility", "Common Misinterpretations": "Treating distributional assumption adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Distributional Assumption Adequacy", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The extent to which distributional assumption is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses distributional assumption adequacy using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in distributional assumption adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
445
79608c041c0babc1613cca804b9e22a2fcbadbeabfda7d7b0bf2336528db7db1
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 306
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Intermediate Precision; Analytical Sensitivity; Analytical Specificity", "Common Misinterpretations": "Treating reproducibility of measurement as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Reproducibility of Measurement", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "The degree to which reproducibility of measurement yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses reproducibility of measurement using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reproducibility of measurement can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
306
83e10528c5f3255388bb807d6bfc99c4e2c93d4f418421fcd184c242521982d5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 485
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Outcome Ascertainment Validity; Differential Follow-up Risk", "Common Misinterpretations": "Treating follow-up completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Follow-up Completeness", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The extent to which all scientifically necessary components of follow-up are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses follow-up completeness using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in follow-up completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
485
a99fe581309ce29c4e888ffbd9a1cba6c99a6948f2a8e99e23714e0f40920b87
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 404
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Interlaboratory Reproducibility; Method Reproducibility; Result Reproducibility", "Common Misinterpretations": "Treating intralaboratory repeatability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Intralaboratory Repeatability", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The degree to which intralaboratory repeatability yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses intralaboratory repeatability using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in intralaboratory repeatability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
404
44bfb4ae3b82bb963a97cc8f3dfed3bd610b7903405993c8f3031931f746d1ce
null value
Scientifically meaningful null value.
BEMO:3100012
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 259
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Batch-Effect Control; Relatedness Control; Differential Expression Robustness", "Common Misinterpretations": "Treating population stratification control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Population Stratification Control", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of population stratification control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses population stratification control using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in population stratification control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
259
c82d16c5b00c0127f5b02c510a48d677f57d7580cba4dcb68dce1141e1686c1d
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 150
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Sufficiency; Evidence Coverage; Evidence Freshness", "Common Misinterpretations": "Treating evidence completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Evidence Completeness", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The extent to which all scientifically necessary components of evidence are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses evidence completeness using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
150
263bd485141b0cc75ee285fb1989c667b83eb378666886a960ed5bd2a610f232
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 345
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Organ-Specific Toxicity Evidence; Carcinogenicity Evidence Strength; Reproductive Toxicity Evidence Strength", "Common Misinterpretations": "Treating genotoxicity evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Genotoxicity Evidence Strength", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The magnitude and credibility of independent evidence supporting genotoxicity evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses genotoxicity evidence strength using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in genotoxicity evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
345
e40592e267890e51ad35407c7dc2af7775fdac9174afe2f08486d28a5e2901b8
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 185
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Genetic Background Control; Environmental Standardization; Intervention Fidelity in Animal Studies", "Common Misinterpretations": "Treating housing and husbandry control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Housing and Husbandry Control", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of housing and husbandry control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses housing and husbandry control using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in housing and husbandry control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
185
d4c26e56573de7f55de0caafd1e2880580fe66b8b4d599787ebf0b2bb3808344
why it matters
Preserves the source rationale.
BEMO:3200003
quality control requirements
Quality-control requirements for computation.
BEMO:3100020
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 133
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Incorporation Bias Risk; Test-Timing Appropriateness; Comparative Test Accuracy", "Common Misinterpretations": "Treating patient flow integrity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Patient Flow Integrity", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of patient flow integrity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses patient flow integrity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in patient flow integrity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
133
67544e05c38d184f5e9fa9830136b1408e7835e115b1fa9c54147ad65e14b99c
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 310
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Cross-Omics Replication; Latent-Factor Stability; Network Reconstruction Robustness", "Common Misinterpretations": "Treating cross-layer directional concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Cross-Layer Directional Concordance", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The degree of agreement in cross-layer directional across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses cross-layer directional concordance using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cross-layer directional concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
310
89f675805a18a02e8214d03853d6a7cdcea2d29079ee839babae7605cea7a581
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 143
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Partial Area Under the Receiver Operating Characteristic Curve; Reference Standard Validity; Index-Test Blinding", "Common Misinterpretations": "Treating threshold validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Threshold Validity", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "The degree to which threshold supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses threshold validity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in threshold validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
143
c27bfe6217acf4b748aeb85aa4e9c4176c8256e12b7c61ff84a201b6fd374c82
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 361
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Toxicokinetic Concordance; Human-Relevance of Toxicological Evidence; Mixture Interaction Assessment", "Common Misinterpretations": "Treating species extrapolation validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Species Extrapolation Validity", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "The degree to which species extrapolation supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses species extrapolation validity using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in species extrapolation validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
361
f3eb077ab3d21773df18d06eb29c8a11272ed7cbfe8f22da0ffedd8a171bc393
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 366
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Pathway Enrichment Robustness", "Common Misinterpretations": "Treating cross-omics concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Developing", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Cross-Omics Concordance", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "The degree of agreement in cross-omics across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses cross-omics concordance using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cross-omics concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
366
447ad3207c1d5333015f6a5189727d810a3a517e4d6a4351f8dc28acbe93a5f0
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 121
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Time-Dependent Discrimination", "Common Misinterpretations": "Treating competing-risk model validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Competing-Risk Model Validity", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "The probability or degree that competing-risk model validity introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses competing-risk model validity using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in competing-risk model validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
121
6f322698057cc9561b413498181b8f2503720420a935d144d2173ae5298a3b2f
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 370
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Protein Identification Confidence; Peptide-Spectrum Match Quality; Protein Inference Reliability", "Common Misinterpretations": "Treating false discovery rate control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "False Discovery Rate Control", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of false discovery rate control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses false discovery rate control using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in false discovery rate control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
370
4059adb92dbe8fea43e57c309bf3c42531450ae52d814c6f79932af37f2f97a2
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 373
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Missing-Value Burden; Retention-Time Stability; Mass Accuracy", "Common Misinterpretations": "Treating ion suppression assessment as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Ion Suppression Assessment", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of ion suppression assessment.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses ion suppression assessment using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in ion suppression assessment can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
373
c108ebf1ca903ed7e96063f24ecaf39723dd5e925f97d641601f89fd859897b9
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 314
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Perturbation Prediction Accuracy", "Common Misinterpretations": "Treating emergent-property reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Specialized / infrequent", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Developing", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Emergent-Property Reproducibility", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The degree to which emergent-property reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses emergent-property reproducibility using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in emergent-property reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
314
e554ab3ba02f294f767076d372c3967762b850eccc356bf4d7c243480641cd98
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 400
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Computational Reproducibility; Interlaboratory Reproducibility; Intralaboratory Repeatability", "Common Misinterpretations": "Treating experimental reproducibility as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Experimental Reproducibility", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "The degree to which experimental reproducibility yields concordant results under the specified repeated-analysis or repeated-measurement conditions.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses experimental reproducibility using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in experimental reproducibility can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
400
6c9d1f649c85d73c0009300283cb79e817b0ca9b1d53444f1a4aa4b5a2e6d8ae
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 63
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Safety Biomarker Validity; Response Biomarker Validity; Surrogate Endpoint Validity", "Common Misinterpretations": "Treating susceptibility/risk biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Susceptibility/Risk Biomarker Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The probability or degree that susceptibility/risk biomarker validity introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses susceptibility/risk biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in susceptibility/risk biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
63
596927f557190db71a0b686158bde6134cafa9cdb9c1876784270d91a80452f5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 172
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Small-Study Effects; Between-Study Variance; Prediction Interval Adequacy", "Common Misinterpretations": "Treating study heterogeneity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Study Heterogeneity", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of study heterogeneity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Forest plots; heterogeneity statistics; tau-squared; prediction intervals; funnel plots; regression or selection models; sensitivity analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses study heterogeneity using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in study heterogeneity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
172
93fce0e52f596e9f5e6b6636a82da697803846e686e6144faa894115a270196a
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 420
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Missing-Data Reporting Completeness; Conflict-of-Interest Transparency; Data-Sharing Transparency", "Common Misinterpretations": "Treating funding-source transparency as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Funding-Source Transparency", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of funding-source transparency.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses funding-source transparency using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in funding-source transparency can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
420
35377c10ba740cabcf147ce4b6ab00e3a9c4087ca2ca810c337f7fe08223e6f9
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 443
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Discrimination of Statistical Predictions; Clinical Relevance of Effect; Fragility Index", "Common Misinterpretations": "Treating decision-curve net benefit as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Decision-Curve Net Benefit", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of decision-curve net benefit.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses decision-curve net benefit using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in decision-curve net benefit can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
443
7e96f3c494e389efab04f0ca564863ebd840cadeaf424446d3f94ddb483ed76c
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 86
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Preservation Adequacy; Freeze–Thaw Burden; Transport Condition Integrity", "Common Misinterpretations": "Treating storage temperature control as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Storage Temperature Control", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of storage temperature control.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses storage temperature control using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in storage temperature control can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
86
b39a92acfa55bd840331823c88476d9d65ca0179f68f0f37ed8d7677a63d2193
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 58
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Content Validity; Prognostic Biomarker Validity; Diagnostic Biomarker Validity", "Common Misinterpretations": "Treating predictive biomarker validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Predictive Biomarker Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which predictive biomarker supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses predictive biomarker validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in predictive biomarker validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
58
c5c7cbb5b7738940f61731df207186f05f8ae71ecab7ca3c5588c823af8d00eb
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 323
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Network Edge Confidence; Module Stability; Pathway Enrichment Consistency", "Common Misinterpretations": "Treating network node confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Network Node Confidence", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "The justified degree of certainty assigned to network node given the quantity, quality, consistency, and limitations of supporting evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses network node confidence using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in network node confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
323
79eaf0a8e6c6fc3012d48e379c545c17d89aecf6de033d40ae6399ec89441a3b
license status
License status of the record or ontology.
BEMO:3200017
output datatype
Expected output datatype.
BEMO:3100008
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 430
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Prospective Registration; Outcome Definition Completeness; Intervention Description Completeness", "Common Misinterpretations": "Treating prespecified analysis adherence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Prespecified Analysis Adherence", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of prespecified analysis adherence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses prespecified analysis adherence using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in prespecified analysis adherence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
430
fe370510831c4feb18d187a9f8348ee0b93ff3174da9ddceea906467255bbd60
Proportion Scale
Controlled BEMO ScaleType value: ProportionScale.
BEMO:4000004
ProportionScale
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 103
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Assumption Sensitivity", "Common Misinterpretations": "Treating falsification test support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Falsification Test Support", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The magnitude and credibility of independent evidence supporting falsification test.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses falsification test support using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in falsification test support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
103
280760fd65dbda26ef711b182bed36c1b7e3679af41be52e50dfca26f92b3400
metric computation specification
An information content entity specifying inputs, outputs, formula status, uncertainty, thresholds, missing-data policy, and validation requirements for computing a metric.
bemo
BEMO:0000100
Candidate
machine formula expression
Serialized machine-readable computation expression.
BEMO:3100005
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 104
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Time-Varying Confounding Control; Reverse-Causation Risk; Selection-on-Survival Bias Risk", "Common Misinterpretations": "Treating immortal-time bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Immortal-Time Bias Risk", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The probability or degree that immortal-time bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses immortal-time bias risk using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in immortal-time bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
104
21a510319cb99a940f9fc99111522f80932b8a874f1ea984fb332f8f218483c8
profile version
Version of a computation profile.
BEMO:3100002
Genetics and Variant Evidence metric
Category of biomedical evidence metrics concerned with genetics and variant evidence.
bemo
BEMO:1100009
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 463
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Bayes Factor Evidence; Prior Sensitivity; Equivalence Margin Validity", "Common Misinterpretations": "Treating posterior probability strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Posterior Probability Strength", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "The magnitude and credibility of independent evidence supporting posterior probability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses posterior probability strength using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in posterior probability strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
463
8a96194c8b8e65d04df266b282704e480656791b819df12c8c4924c984b02272
source record hash
SHA-256 hash of the canonical source record.
BEMO:3100023
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 174
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Animal Model Face Validity; Animal Model Predictive Validity; Species Appropriateness", "Common Misinterpretations": "Treating animal model construct validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Animal Model Construct Validity", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The degree to which animal model construct supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses animal model construct validity using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in animal model construct validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
174
f8c928bcc279a4856d6377104c402fb4a00d102300c3c3fa14181f174b53f974
source frameworks
Preserves the source related frameworks.
BEMO:3200010
calibration requirement
Calibration requirement for the output.
BEMO:3100021
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 159
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Freshness; Evidence Triangulation Strength; Counterevidence Strength", "Common Misinterpretations": "Treating evidence robustness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Evidence Robustness", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence robustness.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses evidence robustness using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence robustness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
159
59315ce160586081fc3fd71f96646716829d26fde9e0a21aa0afec923b5083d6
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 122
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Diagnostic accuracy, screening, prognostic-factor, and prediction-model studies", "Category": "Diagnostic and Prognostic Evidence", "Closely Related Metrics": "Negative Likelihood Ratio; Area Under the Receiver Operating Characteristic Curve; Partial Area Under the Receiver Operating Characteristic Curve", "Common Misinterpretations": "Treating diagnostic odds ratio as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Diagnostic Odds Ratio", "References or Origin": "https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "QUADAS-2; STARD; TRIPOD; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of diagnostic odds ratio.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Two-by-two tables; binomial confidence intervals; hierarchical diagnostic meta-analysis; threshold and prevalence analyses.", "Units or Scale (if applicable)": "Ratio scale; null typically 1", "What It Measures": "Assesses diagnostic odds ratio using evidence appropriate to diagnostic and prognostic evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in diagnostic odds ratio can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
122
a903f03627bce43e756f3c1d10c27c552a53bce0346a40196ea1cd72afe63de0
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 82
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Fixation Adequacy; Storage Temperature Control; Freeze–Thaw Burden", "Common Misinterpretations": "Treating preservation adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Preservation Adequacy", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "The extent to which preservation is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses preservation adequacy using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in preservation adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
82
de3923232eba7ea0cab4f8eddc6a1539eac9f06416d7c31e6d2ce8fbf8492ba5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 48
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Convergent Validity; Criterion Validity; Construct Validity", "Common Misinterpretations": "Treating discriminant validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Discriminant Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which discriminant supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses discriminant validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in discriminant validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
48
c0e22c7d45b8b17a4c9f1de5c7c7bb4e085d98eb667da8d43e452ea31955900b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 99
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized and observational etiologic studies, natural experiments, target-trial emulations", "Category": "Causal Inference", "Closely Related Metrics": "Negative-Control Validation; Dechallenge–Rechallenge Support; Mediation Evidence Strength", "Common Misinterpretations": "Treating dose–response support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Dose–Response Support", "References or Origin": "https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.gradeworkinggroup.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "ROBINS-I; ROBINS-E; GRADE; STROBE; ICH E9", "Scientific Definition": "The magnitude and credibility of independent evidence supporting dose–response.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Method- and analyte-specific physical units", "What It Measures": "Assesses dose–response support using evidence appropriate to causal inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in dose–response support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
99
75f662d9e2ee059c317733c8704b2f0f1f234146f1b2920ed572ec892b5a2d41
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 437
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Deviations-from-Protocol Transparency", "Common Misinterpretations": "Treating reproducibility information completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Reproducibility Information Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of reproducibility information are present, documented, and evaluable.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses reproducibility information completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in reproducibility information completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
437
3aaeec4ac26a2a3b26b615c686a2f19768e2153bd9e1e676bdcd1f2ba2812fc8
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 411
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All experimental, computational, clinical, and omics studies", "Category": "Reproducibility and Replication", "Closely Related Metrics": "Random-Seed Stability; Multiverse Analysis Robustness; Specification-Curve Robustness", "Common Misinterpretations": "Treating researcher-degrees-of-freedom sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Researcher-Degrees-of-Freedom Sensitivity", "References or Origin": "https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.psidev.info/miape | https://arriveguidelines.org/arrive-guidelines", "Related Frameworks": "PRISMA; CONSORT; MIAME; MINSEQE; MIAPE; ARRIVE 2.0", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of researcher-degrees-of-freedom sensitivity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses researcher-degrees-of-freedom sensitivity using evidence appropriate to reproducibility and replication, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in researcher-degrees-of-freedom sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
411
c196ee0e5b3c49cffc78d559a15b400e1dec74f4854200e7fbd341eef60a59c4
scale type
A controlled concept describing the scale used by a metric output.
bemo
BEMO:0000400
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 213
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Generalizability; Population Applicability; Intervention Applicability", "Common Misinterpretations": "Treating setting applicability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Setting Applicability", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of setting applicability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses setting applicability using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in setting applicability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
213
323f595d53815c626b380456c55c73f7ff5e2e10beb1e0ffc44edb8067c37e86
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 151
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Strength; Evidence Stability; Evidence Consistency", "Common Misinterpretations": "Treating evidence confidence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Confidence", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "The justified degree of certainty assigned to evidence given the quantity, quality, consistency, and limitations of supporting evidence.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses evidence confidence using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence confidence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
151
a6abc5b380d07d44d4ec1e48b63cf0e3bf7299ab7264d6edcc898627f3bd0952
has directionality
Connects a computation specification to interpretation directionality.
BEMO:3000011
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 337
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Lowest-Observed-Adverse-Effect Level Robustness; Toxicological Mode-of-Action Support; Adverse Outcome Pathway Support", "Common Misinterpretations": "Treating benchmark dose reliability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Benchmark Dose Reliability", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of benchmark dose reliability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Method- and analyte-specific physical units", "What It Measures": "Assesses benchmark dose reliability using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in benchmark dose reliability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
337
43d4a4ce7127bc52167a879f611c41413072f25e775dbe7c9446cf4028783b10
assessment timestamp
Time at which an assessment was generated.
BEMO:3100027
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 284
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Laboratory-developed tests, molecular assays, imaging, pathology, biomarker studies", "Category": "Measurement and Assay Analytical Validity", "Closely Related Metrics": "Interference Susceptibility; Carryover; Hook Effect Risk", "Common Misinterpretations": "Treating cross-reactivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Cross-Reactivity", "References or Origin": "https://www.fda.gov/media/119271/download | https://clsi.org/standards/products/method-evaluation/ | https://www.iso.org/standard/76677.html | https://pubmed.ncbi.nlm.nih.gov/19246619/", "Related Frameworks": "FDA Biomarker; CLSI; ISO 15189; MIQE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cross-reactivity.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses cross-reactivity using evidence appropriate to measurement and assay analytical validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cross-reactivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
284
1de0ffae2d83059ba1d111ea3bf309f20078d963ac7b7f912a42e8b4e10e5e32
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 200
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Clinical, epidemiologic, diagnostic, translational, and population studies", "Category": "External Validity and Applicability", "Closely Related Metrics": "Care-Pathway Independence; Ecological Validity; Real-World Evidence Alignment", "Common Misinterpretations": "Treating context sensitivity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified threshold; valid reference standard; complete 2×2 classification; confidence intervals; spectrum and prevalence assessment.", "Metric": "Context Sensitivity", "References or Origin": "https://www.gradeworkinggroup.org/ | https://pubmed.ncbi.nlm.nih.gov/22007046/ | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/", "Related Frameworks": "GRADE; QUADAS-2; CONSORT; STROBE", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of context sensitivity.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses context sensitivity using evidence appropriate to external validity and applicability, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in context sensitivity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
200
84e49bfba34c254503c96b181528bf2beab33e58bfa3d6d711b6afaa740ea566
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 483
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Randomized trials, nonrandomized intervention studies, cohort, case-control, cross-sectional studies", "Category": "Study Design and Internal Validity", "Closely Related Metrics": "Cluster Recruitment Bias Risk; Adherence Integrity; Exposure Classification Validity", "Common Misinterpretations": "Treating early stopping bias risk as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Prespecified signaling questions; direction and likely magnitude of distortion; domain-level and overall judgment; sensitivity to plausible bias.", "Metric": "Early Stopping Bias Risk", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.riskofbias.info/welcome/rob-2-0-tool | https://www.riskofbias.info/welcome/home/current-version-of-robins-i | https://www.riskofbias.info/welcome/robins-e-tool | https://www.nhlbi.nih.gov/health-topics/study-quality-assessment-tools | https://jbi.global/critical-appraisal-tools", "Related Frameworks": "CONSORT; STROBE; RoB 2; ROBINS-I; ROBINS-E; NIH Quality Tools; JBI", "Scientific Definition": "The probability or degree that early stopping bias introduces systematic distortion into a biomedical estimate or conclusion.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses early stopping bias risk using evidence appropriate to study design and internal validity, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in early stopping bias risk can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
483
8ac86a4a6b6a886dad127939918a466b8c70bd2b6b13b03a987aab048ac4c7c3
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 440
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All quantitative biomedical studies", "Category": "Statistical Validity and Inference", "Closely Related Metrics": "Measurement Error Correction; Discrimination of Statistical Predictions; Decision-Curve Net Benefit", "Common Misinterpretations": "Treating calibration of statistical predictions as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Calibration of Statistical Predictions", "References or Origin": "https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.tripod-statement.org/ | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/ | https://database.ich.org/sites/default/files/E9_Guideline.pdf", "Related Frameworks": "CONSORT; STROBE; TRIPOD; REMARK; ICH E9", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of calibration of statistical predictions.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses calibration of statistical predictions using evidence appropriate to statistical validity and inference, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in calibration of statistical predictions can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
440
1d5cbda1b0dc5b30aa8339e6219d08168170bd684f74bf0679b26dd4ece9ead4
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 338
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Pharmacology, pharmacokinetics, toxicology, safety and mode-of-action studies", "Category": "Pharmacology and Toxicology", "Closely Related Metrics": "Pharmacodynamic Adequacy; Dose Proportionality; Time–Concentration Profile Adequacy", "Common Misinterpretations": "Treating bioavailability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Bioavailability", "References or Origin": "https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.fda.gov/media/119271/download | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline", "Related Frameworks": "OECD; OHAT; FDA Biomarker; EMA E16", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of bioavailability.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses bioavailability using evidence appropriate to pharmacology and toxicology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in bioavailability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
338
1dbec5a00f53d17cad5a3a88dd4dfde9ceab9bbbe33a898c76c7086698d1a0a2
evidence object
An entity or information artifact evaluated by a biomedical evidence metric, such as a claim, study, dataset, result, hypothesis, assay, or body of evidence.
bemo
BEMO:0000201
Candidate
derived from metric
Represents a curated derivation relation between metrics.
BEMO:3000017
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 265
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Sample Identity Concordance; Variant Call Quality; Genotype Quality", "Common Misinterpretations": "Treating sex concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Sex Concordance", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "The degree of agreement in sex across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses sex concordance using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in sex concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
265
897b2797d8e6449a88906be71f955c7726607e3b28833aef8661c12a890476a4
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 167
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Evidence Quality; Evidence Strength", "Common Misinterpretations": "Treating overall evidence certainty as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Overall Evidence Certainty", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of overall evidence certainty.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses overall evidence certainty using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in overall evidence certainty can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
167
31b48110148a1ebdd1379016bfff93898087ec61573fac3b850ccada3e1f7ae5
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 390
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mass spectrometry, affinity proteomics, metabolomics, lipidomics studies", "Category": "Proteomics and Metabolomics", "Closely Related Metrics": "Ion Suppression Assessment; Mass Accuracy; Isotope Pattern Fidelity", "Common Misinterpretations": "Treating retention-time stability as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Independent repeats; between-run/site/analyst variability; concordance of effect direction and magnitude; predefined reproducibility thresholds.", "Metric": "Retention-Time Stability", "References or Origin": "https://www.psidev.info/miape | https://www.psidev.info/ | https://www.metabolomics-msi.org/", "Related Frameworks": "MIAPE; HUPO PSI; Metabolomics Standards", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of retention-time stability.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Target-decoy analysis; spectral scoring; reference standards; replicate injections; retention-time and mass-error monitoring; orthogonal confirmation.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses retention-time stability using evidence appropriate to proteomics and metabolomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in retention-time stability can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
390
7ed489b83705dc957b891fd6483e7b54ca4c70547f903f7dfa9e53d0ca59fa58
External Validity and Applicability metric
Category of biomedical evidence metrics concerned with external validity and applicability.
bemo
BEMO:1100008
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 221
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "RNA Evidence Strength; Penetrance Evidence; Expressivity Consistency", "Common Misinterpretations": "Treating co-segregation likelihood as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Co-segregation Likelihood", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of co-segregation likelihood.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses co-segregation likelihood using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in co-segregation likelihood can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
221
965b0273f9a305e64b50437dcad85ab39818fad51d561d0ac4da1daaf7d204f2
Genetics, Omics, and Systems Biology metric
Metrics assessing genetic, genomic, transcriptomic, proteomic, metabolomic, multi-omic, and systems-level evidence.
bemo
BEMO:1000007
Candidate
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 429
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All biomedical study reports and data releases", "Category": "Research Transparency and Reporting Completeness", "Closely Related Metrics": "Recruitment Reporting Completeness; Harms Reporting Completeness; Statistical Methods Reporting Completeness", "Common Misinterpretations": "Treating participant flow completeness as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Study report / dataset / evidence package", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Required elements present; traceable provenance; unambiguous definitions; accessible underlying data/materials; documented deviations.", "Metric": "Participant Flow Completeness", "References or Origin": "https://www.equator-network.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.consort-spirit.org/ | https://www.equator-network.org/reporting-guidelines/strobe/ | https://www.equator-network.org/reporting-guidelines/stard/ | https://www.tripod-statement.org/ | https://www.care-statement.org/ | https://www.equator-network.org/reporting-guidelines/cheers/", "Related Frameworks": "EQUATOR; PRISMA; CONSORT; STROBE; STARD; TRIPOD; CARE; CHEERS", "Scientific Definition": "The extent to which all scientifically necessary components of participant flow are present, documented, and evaluable.", "Scientific Importance (1–10)": "9", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Proportion or percentage (0–1 or 0–100%)", "What It Measures": "Assesses participant flow completeness using evidence appropriate to research transparency and reporting completeness, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in participant flow completeness can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
429
c3caaa14ecea68f7075a25ef235cdae89d1a9b9b86ed2b5a30dab5cb5b8bf20d
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 238
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Mendelian disease, cancer genetics, association, segregation, and functional studies", "Category": "Genetics and Variant Evidence", "Closely Related Metrics": "Splicing Evidence Strength; Co-segregation Likelihood; Penetrance Evidence", "Common Misinterpretations": "Treating rna evidence strength as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "RNA Evidence Strength", "References or Origin": "https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://pubmed.ncbi.nlm.nih.gov/25741868/ | https://www.equator-network.org/reporting-guidelines/strega/ | https://geneontology.org/docs/guide-go-evidence-codes/", "Related Frameworks": "ClinGen; ACMG AMP; STREGA; Gene Ontology", "Scientific Definition": "The magnitude and credibility of independent evidence supporting rna evidence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses rna evidence strength using evidence appropriate to genetics and variant evidence, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in rna evidence strength can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
238
8e776a34bb92ed3d093c7eb03481f1c61ce724a0995235355586c28e53418e35
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 54
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Biomarker development, qualification, endpoint and surrogate validation studies", "Category": "Biomarker and Endpoint Validation", "Closely Related Metrics": "Endpoint Responsiveness", "Common Misinterpretations": "Treating minimal clinically important difference validity as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Common within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Mature", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Minimal Clinically Important Difference Validity", "References or Origin": "https://www.fda.gov/media/119271/download | https://www.ncbi.nlm.nih.gov/books/NBK326791/ | https://www.ema.europa.eu/en/ich-e16-genomic-biomarkers-related-drug-response-context-structure-format-qualification-submissions-scientific-guideline | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "FDA Biomarker; BEST; EMA E16; REMARK", "Scientific Definition": "The degree to which minimal clinically important difference supports the intended scientific interpretation without material systematic error.", "Scientific Importance (1–10)": "10", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses minimal clinically important difference validity using evidence appropriate to biomarker and endpoint validation, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in minimal clinically important difference validity can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
54
604e83a95f413c6c8ddc8abe7ad03f24c441ede8e65d506baef0410866c23fa7
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 31
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Network Context Support", "Common Misinterpretations": "Treating systems-level emergence support as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Systems-Level Emergence Support", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The magnitude and credibility of independent evidence supporting systems-level emergence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses systems-level emergence support using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in systems-level emergence support can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
31
07df8f6070c405006e36806e8ccb515126dc2c7269ea9cf3934662d28e9264c6
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 158
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Systematic reviews, meta-analyses, evidence profiles, guidelines", "Category": "Evidence Certainty and Synthesis", "Closely Related Metrics": "Overall Evidence Certainty; Evidence Strength; Evidence Confidence", "Common Misinterpretations": "Treating evidence quality as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Body of evidence / synthesis / outcome", "Frequency of Use": "Common", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Evidence Quality", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.prisma-statement.org/prisma-2020-checklist | https://www.bmj.com/content/358/bmj.j4008 | https://www.riskofbias.info/", "Related Frameworks": "GRADE; PRISMA; AMSTAR 2; RoB", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of evidence quality.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses evidence quality using evidence appropriate to evidence certainty and synthesis, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in evidence quality can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
158
f4e08b1965ce7d1832ff4c5191f1e7298781d9bec3cb9a96c09198d6faaa1339
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 80
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "All studies using human or animal biospecimens", "Category": "Biospecimen and Preanalytical Quality", "Closely Related Metrics": "Cellularity Adequacy; Hemolysis Burden; Lipemia Burden", "Common Misinterpretations": "Treating necrosis burden as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Specimen / assay / run / laboratory / study", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Necrosis Burden", "References or Origin": "https://www.equator-network.org/reporting-guidelines/brisq/ | https://www.iso.org/standard/76677.html | https://www.equator-network.org/reporting-guidelines/reporting-recommendations-for-tumor-marker-prognostic-studies-remark/", "Related Frameworks": "BRISQ; ISO 15189; REMARK", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of necrosis burden.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal risk rating or quantitative percentage/probability; lower is generally better", "What It Measures": "Assesses necrosis burden using evidence appropriate to biospecimen and preanalytical quality, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in necrosis burden can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
80
33466881e023e8dc75d5f189be9da266f95f3bf8283fa46f9f4cff10e5af6f95
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 311
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Integrated omics, networks, pathways, mechanistic and dynamic systems models", "Category": "Multi-omics and Systems Biology", "Closely Related Metrics": "Cross-Omics Replication; Cross-Layer Directional Concordance", "Common Misinterpretations": "Treating cross-omics integration coherence as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Dataset / model / pathway / network / evidence body", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Developing", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Cross-Omics Integration Coherence", "References or Origin": "https://geneontology.org/docs/guide-go-evidence-codes/ | https://reactome.org/ | https://www.uniprot.org/help/evidences | https://www.ga4gh.org/", "Related Frameworks": "Gene Ontology; Reactome; UniProt; GA4GH", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of cross-omics integration coherence.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses cross-omics integration coherence using evidence appropriate to multi-omics and systems biology, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in cross-omics integration coherence can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
311
c10b406fc4ce6db5663280fc4b358cc062cc17f0ce7a593af68a762307815cdc
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 244
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Whole-genome, exome, panel, bulk/single-cell/spatial transcriptomic studies", "Category": "Genomics and Transcriptomics", "Closely Related Metrics": "Genotype Quality; Strand Bias; Reference Bias", "Common Misinterpretations": "Treating allelic balance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Allelic Balance", "References or Origin": "https://www.fged.org/projects/miame | https://www.fged.org/projects/minseqe/ | https://www.equator-network.org/reporting-guidelines/strobe-me/ | https://www.ga4gh.org/ | https://www.humancellatlas.org/", "Related Frameworks": "MIAME; MINSEQE; STROBE-ME; GA4GH; HCA", "Scientific Definition": "A canonical biomedical evidence metric quantifying the scientific credibility, reliability, relevance, or interpretability of allelic balance.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "ClinGen/ACMG evidence scoring; pedigree analysis; population databases; case-control data; functional assays; expert-panel review.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses allelic balance using evidence appropriate to genomics and transcriptomics, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in allelic balance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
244
0857b2292263da43e4b892ac8b30335b2a7fd8730d0baae8269f0919f1da549b
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 178
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "In vitro, ex vivo, organoid, animal, and preclinical experiments", "Category": "Experimental Biology and Animal Research", "Closely Related Metrics": "Experimental Unit Validity; Technical Replicate Adequacy; Sample Size Justification", "Common Misinterpretations": "Treating biological replicate adequacy as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Clear intended use or estimand; accepted reference or criterion; prespecified acceptance thresholds; independent validation; performance across relevant conditions.", "Metric": "Biological Replicate Adequacy", "References or Origin": "https://arriveguidelines.org/arrive-guidelines | https://www.radboudumc.nl/en/research/departments/health-evidence/systematic-review-center-for-laboratory-animal-experimentation | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "ARRIVE 2.0; SYRCLE; OECD", "Scientific Definition": "The extent to which biological replicate is sufficient and fit for the stated biomedical inference.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Ordinal rubric, domain judgment, or normalized score", "What It Measures": "Assesses biological replicate adequacy using evidence appropriate to experimental biology and animal research, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in biological replicate adequacy can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
178
74976e4cf57f53c444f4721c04fc26fb7298cad4ac743db45feeb9a0be34587a
maximum value
Maximum allowed or expected numeric value.
BEMO:3100011
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx row 22
Pure_Biomedical_Evidence_Research_Metrics_Part_2(3).xlsx
{"Applicable Study Types": "Molecular, cellular, animal, translational, pharmacologic, and human studies", "Category": "Biological Plausibility and Mechanism", "Closely Related Metrics": "Phenotypic Concordance; Perturbational Validation; Loss-of-Function Validation", "Common Misinterpretations": "Treating molecular-phenotypic concordance as interchangeable with overall study quality, or interpreting a favorable value as proof that all other bias and validity domains are satisfactory.", "Evidence Level Where Used": "Result / experiment / study / body of evidence, as applicable", "Frequency of Use": "Established within specialty", "Limitations": "Depends on context, operational definition, thresholds, data quality, and evaluator judgment; it should not be used as a stand-alone summary score without domain-level evidence.", "Maturity of Metric": "Established", "Measurement Criteria": "Prespecified operational definition; appropriate comparator or reference; quantified uncertainty; independent or orthogonal corroboration; sensitivity analyses.", "Metric": "Molecular-Phenotypic Concordance", "References or Origin": "https://www.gradeworkinggroup.org/ | https://www.fda.gov/media/119271/download | https://clinicalgenome.org/curation-activities/gene-disease-validity/ | https://ntp.niehs.nih.gov/whatwestudy/assessments/noncancer/handbook | https://www.oecd.org/chemicalsafety/testing/oecd-guidelines-testing-chemicals-related-documents.htm", "Related Frameworks": "GRADE; FDA Biomarker; ClinGen; OHAT; OECD", "Scientific Definition": "The degree of agreement in molecular-phenotypic across measurements, studies, methods, populations, or biological levels.", "Scientific Importance (1–10)": "8", "Typical Methods of Assessment": "Structured critical appraisal; quantitative estimation with uncertainty; prespecified thresholds; independent replication; sensitivity and subgroup analyses.", "Units or Scale (if applicable)": "Metric-specific continuous, categorical, or ordinal scale", "What It Measures": "Assesses molecular-phenotypic concordance using evidence appropriate to biological plausibility and mechanism, distinguishing random uncertainty from systematic error.", "Why It Matters": "Material weakness in molecular-phenotypic concordance can change the direction, magnitude, certainty, or biological interpretation of the research conclusion."}
22
81a78e88eaa62f6e2bb0ae947aba28fe819150032abcf9264f9205803cc7dfd1
normalization method
Normalization method and version.
BEMO:3100014