Abstract
Composite biological-age measures can summarize information distributed across multiple physiological systems, but a single score can also create false precision when measurement conditions, missing data and model changes are hidden. This paper describes a proposed architecture for the WellSkor Funnel Score: a clinician-supervised framework that ingests standardized biomarker data, computes transparent domain subindices, anchors interpretation to a published phenotypic-age measure and preserves every result as an immutable longitudinal record.
The protocol separates a version-controlled marker registry from the computation engine. Each score records its model version, contributing markers, data coverage and collection timepoint. The initial program is designed to establish technical reproducibility, clinical interpretability and analysis-ready longitudinal data. It is not designed to claim disease diagnosis, treatment response, biological-age reversal or outcome prediction before prospective validation.
Scientific rationale
Aging is heterogeneous across people and physiological systems. Published work on Phenotypic Age (Levine et al., 2018; Liu et al., 2018), multisystem dysregulation (Cohen et al., 2013) and longitudinal pace-of-aging measures (Belsky et al., 2015) supports combining information across domains when the derivation, population and intended use are explicit. Different biological-age measures do not necessarily represent the same construct, making repeated measurement and method consistency essential.
WellSkor therefore treats the score as a structured clinical summary rather than an independent biological truth. The intended question is not simply, ‘What is this person’s biological age?’ It is, ‘Which measured domains are contributing to the present profile, how complete is the evidence, and how does the profile change under comparable conditions?’
Proposed data and model architecture
The proposed marker registry would serve as the scientific configuration. For each biomarker it would store a stable code, display name, expected unit, scoring direction, proposed optimal interval, domain assignment, weight and whether the relationship is monotonic or banded. Ranges and weights should never be silently embedded in application code. A frozen snapshot of the registry would become a named model version, allowing an historical result to be reconstructed exactly.
Patient identity, consent, laboratory panels, individual results, domain scores and composite outputs would be stored as separate but linked records. Each score would be keyed to a patient and timepoint and retain the laboratory source, collection date, contributing-marker list, coverage percentage and model version. Recalculation would create a new record rather than overwrite the prior result.
Deterministic computation and quality control
The proposed pipeline is ordered: validate source units and marker identity; normalize each available marker to a 0–100 domain scale using its declared monotonic or banded rule; compute weighted experimental domain subindices; calculate Levine Phenotypic Age using the exact published coefficients and all required inputs; then derive the experimental composite and biological-age gap under the locked model version. Re-normalizing weights across present markers applies only to the proposed WellSkor domain subindices. It must not be used to substitute for a missing Phenotypic Age input while retaining the published measure’s name. In every case, data coverage is displayed rather than concealed.
Technical verification must precede clinical use. Test cases should include known inputs, boundary values, unit conversions, duplicate panels, missing markers and independent recalculation of the published Phenotypic Age component. Sensitivity analyses should quantify how proposed ranges, weights and missingness rules affect the composite. Acute illness, fasting state, medication changes, laboratory platform and pre-analytic variation must remain visible because they can mimic or obscure biological change.
Longitudinal interpretation
A trajectory is meaningful only when the underlying method is comparable. Consecutive clinical reports should identify the model version and highlight changes in data coverage. Serial composite-index methods likewise emphasize consistent item content and sensitivity testing (Searle et al., 2008). If the registry is updated, investigators should either analyze scores within version or re-score all eligible timepoints under one frozen analysis version while preserving the originally reported values.
The composite should be interpreted beside its component domains and conventional clinical findings. A favorable movement in one index does not establish treatment benefit, and a stable composite can hide offsetting changes. The proposed label ‘resilience’ is an operational hypothesis and must not be equated with validated dynamic recovery-time constructs derived from dense repeated measurements (Pyrkov et al., 2021). The protocol therefore prioritizes direction, uncertainty, component-level explanation and clinically meaningful functional outcomes over a single headline number.
Prospective validation plan
The first implementation phase is a feasibility and reproducibility study, not efficacy validation. It should establish reliable data capture, deterministic recalculation, acceptable missingness, clinician interpretability and repeat-measurement workflows. Early de-identified cases may reveal engineering defects or implausible behavior, but a small convenience sample cannot validate coefficients or support predictive claims.
A later protocol should pre-specify the population, primary outcome, follow-up interval, sample-size rationale and analysis plan before coefficients are fitted. Candidate outcomes may include change in a validated aging measure, physical function or adjudicated clinical events, but the primary endpoint must be chosen in advance. Development should report calibration, uncertainty, missing-data handling and subgroup performance, followed by temporal and external validation. Early live evaluation should document workflow, human factors, failures and version changes in the spirit of DECIDE-AI, while any individual prediction study should follow applicable TRIPOD+AI reporting principles.
Clinical governance, privacy and regulatory boundaries
WellSkor is proposed as clinician-facing decision support. A qualified clinician must be able to inspect the source data, transformation rules, coverage and rationale and must independently evaluate the patient. The score is not intended for autonomous diagnosis, treatment selection or time-critical decisions. Whether a deployed software function qualifies as non-device clinical decision support depends on its actual intended use, claims and functionality and requires formal regulatory review.
Any clinical deployment should use role-based access, authentication, encryption, integrity controls, audit logging and business-associate agreements where required for services handling protected health information. Clinical-care consent should be distinguished from permission for secondary research use when required by the governing protocol or applicable law; research on public attitudes indicates that both consent and the stated purpose of data use influence perceived acceptability (Grande et al., 2014). Research datasets should be de-identified under an approved governance plan, and any withdrawal rights should be implemented as specified by the consent, protocol and applicable law.
Limitations and conclusion
The proposed domain structure, optimal ranges and weights are hypotheses. A composite may obscure clinically important individual abnormalities; missingness may be informative rather than random; and Phenotypic Age was developed for population-level risk estimation, not as a universal surrogate endpoint for an individual intervention. Measurement repeatability does not establish biological validity, and association with an outcome does not establish that changing the score changes that outcome.
The WellSkor Funnel Score should advance only through transparent versioning, reproducible computation, prospective validation and restrained claims. Its near-term contribution is a disciplined longitudinal record that helps clinician and patient discuss measured change. Diagnostic, prognostic or intervention-response claims belong to a later research program and should be made only if adequately powered internal and external validation support them.
