Scoring reference
Scoring Guide & FAQ
A reference for external benchmarking participants. How the score is constructed, what the rubric labels mean in practice, and answers to the questions that arise most when scoring for the first time.
Overview
BioMetre measures how well process engineers can see and use the process data they need. Each product-at-a-site is scored on an 18 × 6 matrix producing a weighted composite score from 0 to 100.
- Rows (WHAT): 18 subdimensions of data across 6 categories — what data is available.
- Columns (HOW): 6 quality dimensions — how well that data is delivered.
Data categories are weighted by how directly they bear on understanding process variability: CQAs and Yield — the key outputs, the Y in Y = F(X) — carry 30% each; Batch, Continuous, Genealogy, and Calculated Features (Input) carry 10% each. The six HOW dimensions are weighted equally (1/6 each).
Composite Score Calculation
Rate each combination
For each data category, score its three subdimensions against the six delivery qualities on the 0/25/50/75/100 scale.
Combine by weight
Delivery quality scores average within each subdimension; subdimensions combine by their configured weights (40/40/20) into a category score; categories combine by their impact weight into the overall BIOMETRE score.
Identify priority improvements
Each gap is ranked by the score improvement it would unlock if raised to 100, directing attention toward the highest-impact changes first.
Formula
composite = SUM over all 18 subdimensions of: (category_weight × subdimension_weight) × MEAN(Fresh, Frictionless, Accessible, Authentic, Standard, Structured)
Scoring in Practice — A Worked Example
The table below shows how DS CQA scores evolve across nine infrastructure maturity stages — from paper-only to self-service — using the canonical rubric labels.
| Stage | Consumption maturity — identical in both cases | Data layer unvalidated | Data layer validated | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Fresh | Frictionless | Accessible | Structured | Standard | Authentic | Avg | Authentic | Avg | |
| 1. Paper only | unavailable | day | inaccessible | no context | local | issues | 0 | issues | 0 |
| 2. Paper to spreadsheet | quarterly | 30 min | data engineer | some context | local | unproven | 21 | unproven | 21 |
| 3. Electronic data capture | weekly | 10 min | data engineer | most context | limited | source only | 42 | source only | 42 |
| 4. LIMS (direct login) | daily | 30 min | few local | most context | moderate | E2E | 58 | E2E | 58 |
| 5. LIMS to data lake (raw) | daily | 10 min | many local | table export | limited | source only | 58 | data layer | 63 |
| 6. Data lake + contextualized pipeline | daily | 10 min | globally | table in lake | scalable | source only | 75 | data layer | 79 |
| 7. + Validated SPC tool (sole viz) | daily | 10 min | few local | table in lake | scalable | source only | 67 | E2E | 75 |
| 8. + Modern SPC replacement (sole viz) | daily | 3 min | many local | table in lake | scalable | source only | 75 | E2E | 83 |
| 9. + Self-service interface | daily | 1 min | globally | table in lake | scalable | source only | 83 | data layer | 88 |
Key observations
Stage 6 (contextualized data pipeline to a data lake) lifts the score by ~17 points, addressing Structured, Accessible, Standard, and Fresh simultaneously.
The nine stages describe consumption maturity. How deep validation reaches is a separate property of the site, which is why Authentic appears twice above. The gap between the two Avg columns is what declaring the data layer validated is worth.
Stage 7 scores below Stage 6 because it is a substitution, not an addition — the SPC tool becomes the sole visualization, so Accessible falls from "globally" to "few local." Widening reach can only raise a score; narrowing it lowers one.
A validated tool on an unvalidated data layer is a house of cards: Authentic scores the weakest link, so Stage 7 stays at "source only" (67) until the layer beneath it is validated (75). Validating the consumption tool first buys nothing.
Scoring Tips
- Score what exists today, not what is planned or aspirational.
- Score the typical consumption path, not the best theoretical access.
- Score the artifact the intended many act on — the audience is the one named in the capability statement, not whoever happens to hold a path today.
- Measure reach against that audience, never against current users. A tool only a data engineer can run is not "accessible" merely because every data engineer can run it.
- When in doubt, score conservatively — it is better to show improvement over time than to start optimistically.
- Each cell is independent — a product can score 100 on Fresh but 0 on Structured for the same subdimension.
- CMO products are scored from the sponsor's perspective — what can the sponsoring company see, not what the CMO has internally.
- Not all subdimensions apply equally — PAT relevance varies by product type, but score what exists rather than marking "not applicable."
Frequently Asked Questions
What does Fresh measure for CQAs and other analytical results?
Freshness measures the latency between when a data point is finalized at source (e.g., result released in LIMS) and when it is available in the enterprise data layer — not the time required to generate the measurement itself. A 14-day stability assay that is pushed to the data lake the instant it is released is "live" (100). Conversely, a rapid in-process test whose results sit in a local system for weeks until manually exported is poor freshness.
How should system reliability and outages affect scoring?
Reliability is scored in Fresh, not Accessible. Fresh reflects the typical delay in getting current data, not the best case: if a system is nominally live (100) but experiences weekly multi-day outages, score Fresh as if data arrives weekly (50), or worse depending on severity. Data-entry lag counts the same way. Accessible asks a separate question — can the intended many reach the tool at all (reach, discovery, permission) — so an outage is counted once, in Fresh, never twice.
How is Authentic scored when the chain is only partly validated?
Authentic scores the weakest link in the chain beneath the artifact you act on: source system → underlying data layer → user consumption layer. You cannot validate downstream of unassured data, so a validated tool sitting on an unvalidated or merely qualified data layer scores "source only" (50) — below an unvalidated tool on a validated data layer, which scores "data layer" (75). Qualified is not assured; it never lifts a rung on its own.
If a validated path exists but nobody uses it, does it count?
No. Biometre scores what people actually receive, not what the organization possesses. Score the chain beneath the artifact the intended many act on. A validated path that no one consumes from does not raise Authentic — it makes the recommendation cheap: move consumption onto it.
What if users export to Excel or JMP to do the real analysis?
If they go there because the tool lacks the functionality they need, then that spreadsheet is the artifact they act on: it becomes the consumption layer, the tool beneath it becomes the data layer, and Authentic is scored accordingly. If they use it for convenience while still acting on the validated tool's numbers, it is not the artifact and nothing is penalized.
How should CMO products be scored?
Score based on what your organization can see, not the CMO's internal capability. If the CMO has excellent internal systems but your organization only receives quarterly PDF reports, the scores should reflect your access.
What about non-applicable subdimensions (e.g., PAT for non-bio products)?
If a subdimension is genuinely not applicable to a product type, score it as 0. The model's weighting (20% for the third subdimension in each category) limits the impact. Do not skip or leave blank.
How should corporate standard SPC / dashboard tools be scored?
Score the current corporate standard tool on its actual usability: are dashboards named clearly? Is data structured and accessible, or do users struggle to find what they need? Mixed scalability and cryptic naming should be reflected in lower Frictionless and Structured scores.
Why are Continuous Data or Features scores often low?
A low score in these categories often reflects a lack of inherent measurements rather than a failure of data pipelines or tools. This points to the need for PAT (Process Analytical Technology) and soft sensors to generate the data in the first place. The model deliberately captures this gap: if the measurement does not exist, the downstream data infrastructure cannot compensate. Improvement requires investment in sensing capability, not just IT systems.
Glossary
| Term | Definition |
|---|---|
| BioMetre | Balanced Information Observability for Manufacturing Excellence — Transformation Readiness Evaluation. The scoring framework. |
| Composite | Weighted average score across all 108 cells (0–100). |
| Grid | The 18 × 6 matrix of scores for one product-at-a-site. |
| Rubric | The descriptive label for a numeric score (e.g., "daily" for Fresh = 75). |
| E2E | End-to-end — validation reaches from the source system through to the artifact the user acts on. |
| Data layer | The lake or repositories of exported source data, plus their pipelines. Can be declared validated independently of the tools built on it. |
| Consumption layer | The interface the user consumes from, including its exports. Which product is the data layer and which the consumption layer is functional, not fixed. |
| Source only | Validation reaches the source system but no further — the data layer is unvalidated or merely qualified. |
| Best Practice | Score 80–100 (green) — Mature across what and how; ready for AI/ML. |
| High maturity | Score 60–79 (light green) — A good base for moving toward AI/ML; improvements needed. |
| Moderate maturity | Score 40–59 (yellow) — Targeted improvements required. |
| Low maturity | Score < 40 (red) — Significant transformation needed. |