Scoring reference

Scoring Guide & FAQ

A reference for external benchmarking participants. How the score is constructed, what the rubric labels mean in practice, and answers to the questions that arise most when scoring for the first time.

Overview

BioMetre measures how well process engineers can see and use the process data they need. Each product-at-a-site is scored on an 18 × 6 matrix producing a weighted composite score from 0 to 100.

  • Rows (WHAT): 18 subdimensions of data across 6 categories — what data is available.
  • Columns (HOW): 6 quality dimensions — how well that data is delivered.

Data categories are weighted by how directly they bear on understanding process variability: CQAs and Yield — the key outputs, the Y in Y = F(X) — carry 30% each; Batch, Continuous, Genealogy, and Calculated Features (Input) carry 10% each. The six HOW dimensions are weighted equally (1/6 each).

Composite Score Calculation

Step 1

Rate each combination

For each data category, score its three subdimensions against the six delivery qualities on the 0/25/50/75/100 scale.

Step 2

Combine by weight

Delivery quality scores average within each subdimension; subdimensions combine by their configured weights (40/40/20) into a category score; categories combine by their impact weight into the overall BIOMETRE score.

Step 3

Identify priority improvements

Each gap is ranked by the score improvement it would unlock if raised to 100, directing attention toward the highest-impact changes first.

Formula

composite = SUM over all 18 subdimensions of:
  (category_weight × subdimension_weight) × MEAN(Fresh, Frictionless, Accessible, Authentic, Standard, Structured)

Scoring in Practice — A Worked Example

The table below shows how DS CQA scores evolve across nine infrastructure maturity stages — from paper-only to self-service — using the canonical rubric labels.

Stage Consumption maturity — identical in both cases Data layer unvalidated Data layer validated
Fresh Frictionless Accessible Structured Standard Authentic Avg Authentic Avg
1. Paper only unavailable day inaccessible no context local issues 0 issues 0
2. Paper to spreadsheet quarterly 30 min data engineer some context local unproven 21 unproven 21
3. Electronic data capture weekly 10 min data engineer most context limited source only 42 source only 42
4. LIMS (direct login) daily 30 min few local most context moderate E2E 58 E2E 58
5. LIMS to data lake (raw) daily 10 min many local table export limited source only 58 data layer 63
6. Data lake + contextualized pipeline daily 10 min globally table in lake scalable source only 75 data layer 79
7. + Validated SPC tool (sole viz) daily 10 min few local table in lake scalable source only 67 E2E 75
8. + Modern SPC replacement (sole viz) daily 3 min many local table in lake scalable source only 75 E2E 83
9. + Self-service interface daily 1 min globally table in lake scalable source only 83 data layer 88

Key observations

1
The data pipeline is the biggest single lever.

Stage 6 (contextualized data pipeline to a data lake) lifts the score by ~17 points, addressing Structured, Accessible, Standard, and Fresh simultaneously.

2
Authentic is not a stage — it is a second axis.

The nine stages describe consumption maturity. How deep validation reaches is a separate property of the site, which is why Authentic appears twice above. The gap between the two Avg columns is what declaring the data layer validated is worth.

3
Adding a tool never reduces a score. Replacing one can.

Stage 7 scores below Stage 6 because it is a substitution, not an addition — the SPC tool becomes the sole visualization, so Accessible falls from "globally" to "few local." Widening reach can only raise a score; narrowing it lowers one.

4
Validate bottom-up, not top-down.

A validated tool on an unvalidated data layer is a house of cards: Authentic scores the weakest link, so Stage 7 stays at "source only" (67) until the layer beneath it is validated (75). Validating the consumption tool first buys nothing.

Scoring Tips

  1. Score what exists today, not what is planned or aspirational.
  2. Score the typical consumption path, not the best theoretical access.
  3. Score the artifact the intended many act on — the audience is the one named in the capability statement, not whoever happens to hold a path today.
  4. Measure reach against that audience, never against current users. A tool only a data engineer can run is not "accessible" merely because every data engineer can run it.
  5. When in doubt, score conservatively — it is better to show improvement over time than to start optimistically.
  6. Each cell is independent — a product can score 100 on Fresh but 0 on Structured for the same subdimension.
  7. CMO products are scored from the sponsor's perspective — what can the sponsoring company see, not what the CMO has internally.
  8. Not all subdimensions apply equally — PAT relevance varies by product type, but score what exists rather than marking "not applicable."

Frequently Asked Questions

What does Fresh measure for CQAs and other analytical results?

Freshness measures the latency between when a data point is finalized at source (e.g., result released in LIMS) and when it is available in the enterprise data layer — not the time required to generate the measurement itself. A 14-day stability assay that is pushed to the data lake the instant it is released is "live" (100). Conversely, a rapid in-process test whose results sit in a local system for weeks until manually exported is poor freshness.

How should system reliability and outages affect scoring?

Reliability is scored in Fresh, not Accessible. Fresh reflects the typical delay in getting current data, not the best case: if a system is nominally live (100) but experiences weekly multi-day outages, score Fresh as if data arrives weekly (50), or worse depending on severity. Data-entry lag counts the same way. Accessible asks a separate question — can the intended many reach the tool at all (reach, discovery, permission) — so an outage is counted once, in Fresh, never twice.

How is Authentic scored when the chain is only partly validated?

Authentic scores the weakest link in the chain beneath the artifact you act on: source system → underlying data layer → user consumption layer. You cannot validate downstream of unassured data, so a validated tool sitting on an unvalidated or merely qualified data layer scores "source only" (50) — below an unvalidated tool on a validated data layer, which scores "data layer" (75). Qualified is not assured; it never lifts a rung on its own.

If a validated path exists but nobody uses it, does it count?

No. Biometre scores what people actually receive, not what the organization possesses. Score the chain beneath the artifact the intended many act on. A validated path that no one consumes from does not raise Authentic — it makes the recommendation cheap: move consumption onto it.

What if users export to Excel or JMP to do the real analysis?

If they go there because the tool lacks the functionality they need, then that spreadsheet is the artifact they act on: it becomes the consumption layer, the tool beneath it becomes the data layer, and Authentic is scored accordingly. If they use it for convenience while still acting on the validated tool's numbers, it is not the artifact and nothing is penalized.

How should CMO products be scored?

Score based on what your organization can see, not the CMO's internal capability. If the CMO has excellent internal systems but your organization only receives quarterly PDF reports, the scores should reflect your access.

What about non-applicable subdimensions (e.g., PAT for non-bio products)?

If a subdimension is genuinely not applicable to a product type, score it as 0. The model's weighting (20% for the third subdimension in each category) limits the impact. Do not skip or leave blank.

How should corporate standard SPC / dashboard tools be scored?

Score the current corporate standard tool on its actual usability: are dashboards named clearly? Is data structured and accessible, or do users struggle to find what they need? Mixed scalability and cryptic naming should be reflected in lower Frictionless and Structured scores.

Why are Continuous Data or Features scores often low?

A low score in these categories often reflects a lack of inherent measurements rather than a failure of data pipelines or tools. This points to the need for PAT (Process Analytical Technology) and soft sensors to generate the data in the first place. The model deliberately captures this gap: if the measurement does not exist, the downstream data infrastructure cannot compensate. Improvement requires investment in sensing capability, not just IT systems.

Glossary

Term Definition
BioMetre Balanced Information Observability for Manufacturing Excellence — Transformation Readiness Evaluation. The scoring framework.
Composite Weighted average score across all 108 cells (0–100).
Grid The 18 × 6 matrix of scores for one product-at-a-site.
Rubric The descriptive label for a numeric score (e.g., "daily" for Fresh = 75).
E2E End-to-end — validation reaches from the source system through to the artifact the user acts on.
Data layer The lake or repositories of exported source data, plus their pipelines. Can be declared validated independently of the tools built on it.
Consumption layer The interface the user consumes from, including its exports. Which product is the data layer and which the consumption layer is functional, not fixed.
Source only Validation reaches the source system but no further — the data layer is unvalidated or merely qualified.
Best Practice Score 80–100 (green) — Mature across what and how; ready for AI/ML.
High maturity Score 60–79 (light green) — A good base for moving toward AI/ML; improvements needed.
Moderate maturity Score 40–59 (yellow) — Targeted improvements required.
Low maturity Score < 40 (red) — Significant transformation needed.
Open the calculator →