# Retrospective benchmark

Measured potency from ChEMBL, scored through the ordinary pipeline. An AUC of 0.5 is a coin flip. Baselines are there because actives in a medicinal-chemistry series tend to be larger than the compounds they were optimised from, so a scorer that merely prefers big molecules can post a respectable AUC while knowing nothing.

## Headline

| Target | Compounds | Active | OncoForge AUC | 95% CI | Best baseline | Best single axis | Verdict |
| --- | --- | --- | --- | --- | --- | --- | --- |
| **EGFR** | 1141 | 805 | 0.701 | 0.670–0.737 | heavy atoms 0.712 | pharmacophore 0.827 | no better than a trivial baseline |
| **PARP1** | 1518 | 1435 | 0.676 | 0.604–0.744 | heavy atoms 0.726 | precedent 0.746 | no better than a trivial baseline |
| **PRMT5** | 255 | 216 | 0.469 | 0.341–0.590 | heavy atoms 0.732 | admet 0.647 | no better than a trivial baseline |

## EGFR

- ChEMBL target `CHEMBL203`, 4604 unique compounds available, 1141 scored on the `holdout` split
- Active at pChEMBL ≥ 7 (805), inactive at ≤ 5 (336); compounds between the thresholds are excluded from the classification metrics

| Ranking | AUC | BEDROC(α=20) | EF@1% | EF@5% | Spearman ρ |
| --- | --- | --- | --- | --- | --- |
| **oncoforge total** | 0.701 | 0.944 | 1.42 | 1.39 | +0.317 |
| _axis: pharmacophore_ | 0.827 | 0.946 | 1.42 | 1.32 | +0.427 |
| _axis: precedent_ | 0.631 | 0.917 | 1.16 | 1.32 | +0.264 |
| _axis: geometry_ | 0.589 | 0.941 | 1.42 | 1.34 | +0.251 |
| _axis: safety_ | 0.578 | 0.742 | 1.03 | 1.07 | +0.119 |
| _axis: learned_ | 0.500 | 0.691 | 0.64 | 0.92 | +0.000 |
| _axis: selectivity_ | 0.500 | 0.691 | 0.64 | 0.92 | +0.000 |
| _axis: properties_ | 0.479 | 0.484 | 0.00 | 0.65 | +0.021 |
| _axis: admet_ | 0.314 | 0.207 | 0.00 | 0.17 | -0.184 |
| _axis: synthesis_ | 0.259 | 0.303 | 0.26 | 0.27 | -0.337 |
| heavy atoms | 0.712 | 0.828 | 1.42 | 1.14 | +0.191 |
| molecular weight | 0.710 | 0.825 | 1.42 | 1.12 | +0.219 |
| aromatic rings | 0.693 | 0.743 | 0.77 | 0.99 | +0.194 |
| cLogP | 0.637 | 0.718 | 1.03 | 0.94 | +0.133 |
| random | 0.514 | 0.743 | 1.03 | 1.14 | +0.024 |
| QED | 0.375 | 0.338 | 0.00 | 0.37 | -0.110 |

**Verdict: no better than a trivial baseline.** AUC 0.701 (95% CI 0.670-0.737) does not exceed 'heavy atoms' at 0.712. Whatever ranking ability is on show here is explained by that single property, so the composite score has not yet earned its complexity. Note that 'axis: pharmacophore' alone reaches 0.827, above both the composite and every baseline: the composite is diluting a signal that one of its own axes already carries.

**Covalent subset** — where the geometry axis makes a claim

- 292 compounds carry a warhead (230 potent)
- Predicted reach as a ranker: AUC 0.833; reach × reactivity: AUC 0.787
- Judged able to reach the residue: 62% of potent versus 15% of weak

**Diagnostics**

- the required pharmacophore separates the classes: 93% of potent versus 42% of weak ligands carry it.

## PARP1

- ChEMBL target `CHEMBL3105`, 4574 unique compounds available, 1518 scored on the `holdout` split
- Active at pChEMBL ≥ 7 (1435), inactive at ≤ 5 (83); compounds between the thresholds are excluded from the classification metrics

| Ranking | AUC | BEDROC(α=20) | EF@1% | EF@5% | Spearman ρ |
| --- | --- | --- | --- | --- | --- |
| **oncoforge total** | 0.676 | 0.985 | 1.06 | 1.06 | +0.037 |
| _axis: precedent_ | 0.746 | 0.997 | 1.06 | 1.06 | +0.466 |
| _axis: properties_ | 0.711 | 0.971 | 1.06 | 1.03 | +0.009 |
| _axis: pharmacophore_ | 0.642 | 0.978 | 1.06 | 1.04 | -0.106 |
| _axis: safety_ | 0.566 | 0.974 | 1.06 | 1.04 | +0.045 |
| _axis: geometry_ | 0.508 | 0.955 | 1.06 | 1.03 | +0.057 |
| _axis: selectivity_ | 0.504 | 0.957 | 0.99 | 1.04 | -0.100 |
| _axis: learned_ | 0.500 | 0.956 | 1.06 | 1.03 | +0.000 |
| _axis: admet_ | 0.458 | 0.891 | 1.06 | 0.97 | +0.036 |
| _axis: synthesis_ | 0.219 | 0.619 | 0.71 | 0.71 | -0.248 |
| heavy atoms | 0.726 | 0.943 | 0.99 | 1.00 | +0.305 |
| molecular weight | 0.713 | 0.934 | 0.92 | 1.02 | +0.284 |
| aromatic rings | 0.636 | 0.931 | 0.99 | 1.02 | +0.158 |
| cLogP | 0.533 | 0.842 | 0.71 | 0.95 | -0.029 |
| QED | 0.511 | 0.930 | 1.06 | 1.00 | -0.124 |
| random | 0.502 | 0.967 | 1.06 | 1.04 | -0.034 |

**Verdict: no better than a trivial baseline.** AUC 0.676 (95% CI 0.604-0.744) does not exceed 'heavy atoms' at 0.726. Whatever ranking ability is on show here is explained by that single property, so the composite score has not yet earned its complexity.

**Covalent subset** — where the geometry axis makes a claim

- 32 compounds carry a warhead (29 potent)
- Judged able to reach the residue: 0% of potent versus 0% of weak
- _every warhead-bearing compound received the same reach score, so no ranking was possible — usually because the target has no reactive residue and these electrophiles are incidental rather than designed_

**Diagnostics**

- the required pharmacophore is detected almost as often in weak ligands (41%) as in potent ones (53%), so its presence carries little information here.
- the pocket-size model expects about 25 heavy atoms, but ligands that actually bind have a median of 30. The shape term is therefore penalising the compounds it should reward; the pocket volume for this target is too small, or the volume-to-heavy-atom constant is.

## PRMT5

- ChEMBL target `CHEMBL1795116`, 1062 unique compounds available, 255 scored on the `holdout` split
- Active at pChEMBL ≥ 7 (216), inactive at ≤ 5 (39); compounds between the thresholds are excluded from the classification metrics

| Ranking | AUC | BEDROC(α=20) | EF@1% | EF@5% | Spearman ρ |
| --- | --- | --- | --- | --- | --- |
| **oncoforge total** | 0.469 | 0.778 | 0.79 | 1.00 | -0.020 |
| _axis: admet_ | 0.647 | 0.904 | 1.18 | 1.00 | +0.230 |
| _axis: properties_ | 0.534 | 0.856 | 1.18 | 1.09 | +0.120 |
| _axis: precedent_ | 0.514 | 0.881 | 1.18 | 1.00 | -0.083 |
| _axis: geometry_ | 0.500 | 0.874 | 1.18 | 1.09 | +0.000 |
| _axis: learned_ | 0.500 | 0.874 | 1.18 | 1.09 | +0.000 |
| _axis: selectivity_ | 0.500 | 0.874 | 1.18 | 1.09 | +0.000 |
| _axis: safety_ | 0.423 | 0.792 | 1.18 | 0.91 | -0.226 |
| _axis: pharmacophore_ | 0.370 | 0.439 | 0.39 | 0.36 | -0.035 |
| _axis: synthesis_ | 0.279 | 0.732 | 1.18 | 1.00 | -0.208 |
| heavy atoms | 0.732 | 0.852 | 1.18 | 0.91 | +0.051 |
| molecular weight | 0.694 | 0.877 | 1.18 | 1.00 | +0.038 |
| aromatic rings | 0.526 | 0.902 | 1.18 | 1.09 | -0.068 |
| cLogP | 0.439 | 0.481 | 0.00 | 0.54 | -0.087 |
| random | 0.420 | 0.704 | 0.39 | 0.91 | +0.001 |
| QED | 0.416 | 0.532 | 0.00 | 0.73 | +0.016 |

**Verdict: no better than a trivial baseline.** AUC 0.469 (95% CI 0.341-0.590) does not exceed 'heavy atoms' at 0.732. Whatever ranking ability is on show here is explained by that single property, so the composite score has not yet earned its complexity.

**Diagnostics**

- the pharmacophore this target requires (cofactor_mimetic) is recognised in only 0% of genuinely potent ligands, so the pharmacophore axis is scoring noise. The substructure vocabulary for this target is missing, not merely imprecise.
- the pocket-size model expects about 26 heavy atoms, but ligands that actually bind have a median of 33. The shape term is therefore penalising the compounds it should reward; the pocket volume for this target is too small, or the volume-to-heavy-atom constant is.

