PLAYER 1 · READY

Chukwudi Eke

Physician · Independent ML researcher

Open to data science and ML roles · remote-friendly

I treat models like patients: pre-register the trial, then report what it found.

Evaluation and calibration for medical-imaging models, world models, and football analytics. I care less about whether a model looks good than whether its confidence can be trusted when the data moves.

CHUKS
Class · Physician-researcher

HIGH SCORES

Robot episodes judged
6,742
CT cases in a locked dataset
6,212
Pre-registered studies
2
Hypotheses refuted and reported
1
01

QUEST LOG

Showing 13 of 13 projects

Judgetron
Shipped
Does a VLM robot-failure judge's confidence survive the move to a new robot?

LoRA lifts in-domain AUROC 0.588 → 0.775 but only 0.650 → 0.690 on a new robot. An auto-accept threshold validated at 1.8% error runs at 26.5% on deployment.

Error on auto-accepted verdictsthreshold set for 5%
Judgetron: error on auto-accepted verdicts versus the 5% target
SettingAuto-accept rateRealised errorVersus target
Same robots (RLBench + WidowX)14.5%1.8%within target
New robot (UR5, never seen)7.9%26.5%5.3 times over
VLMCalibrationConformal
No Safe Tier
Paper
Change only a patient's location. Does the model's differential diagnosis change, and is that safe?

Scale stops models over-calling tropical disease, but under-calling of treatable cosmopolitan disease in a low-income framing rises with scale. Three frontier models under-call at 0.82–1.00. Not scale, not medical fine-tuning, not the frontier: there is no safe tier.

LLM auditDiagnostic biasGlobal health
Prolepsis
Running
Do contrastive world models keep the state that pixel reconstruction throws away?

Pre-registered and hash-locked before a single grid run. 20 runs on a synthetic arcade world, decision rule fixed in advance.

World modelsPre-registrationContrastive
DIRAC
Refuted
A quantum-inspired structured latent for 3D CT, against parameter-matched CNN and ViT baselines.

The advantage was real and reproducible, and ablations traced all of it to one training loss, not the architecture. Hypothesis refuted, and said so.

Medical imagingFoundation modelsAblations
covtoken
Complete
A label-free lesion subspace inside frozen self-supervised medical vision transformers.

Mid-layer geometry localises lesions without labels across CT and ultrasound, MedDINOv3 and DINOv2. Membership pruning beats saliency on small-lesion miss-rate, with a conformal retention certificate and 1.6× fewer FLOPs.

SSLToken pruningConformal
Vigil
Shipped
Pharmacovigilance for clinicians: know before you prescribe, for any drug, in under 60 seconds.

Scans 11M+ biomedical papers and live regulatory sources for safety signals, interactions, patient-specific dosing, pharmacogenomics, and Africa formulary status (NAFDAC · SAHPRA · WHO).

Drug safetyClinical toolsAfrica
MALIT v2
Paper
Malaria screening from thin blood smears that knows when to hand off to a human.

A biologically inspired CNN with learnable Gabor filters and competitive inhibition, trained with a differentiable calibration term (CE + λ·ECE) and confidence-gated escalation.

MicroscopyCalibrationEscalation
Safe clinical fusion
Live
Clinical decision support that is safe by construction, for low-resource settings.

Fuses per-modality signals, escalates to a clinician on contraindications, hash-signs every output for audit, and degrades gracefully when a modality is missing. Runs in the browser, no install.

Decision supportSafetyLow-resource
Fulcrum
Running
A player-agnostic geometric world model for football: one canonical state from tracking, broadcast or freeze-frames.

Finds exploitable space by running persistent homology over a velocity-aware pitch-control field. Validated as a shot precursor at AUC 0.886.

FootballTopologyWorld models
Fulcrum Scout
Live
Recruitment intelligence that scouts by what a player's geometry demonstrates, not just what they produced.

Every claim carries a machine-readable scientific tier so the interface can't over-claim, and signing impact is simulated live through a validated tactical twin.

ScoutingCounterfactualsStreamlit
Stoichima
Shipped
Full-stack match prediction for the Premier League and La Liga.

An XGBoost ensemble, a Dixon-Coles goals model and a Bi-LSTM sequence model behind a FastAPI service, covering outcome, goals, both-teams-to-score, correct-score and handicap markets.

XGBoostFastAPIForecasting
wavelet-grad
Complete
Wavelet denoising inside the optimizer's update step.

WaveletAdam solves noisy XOR 10/10 against Adam's 7/10 at σ = 0.05, with 6× lower final loss. Pure NumPy, 53+ tests.

OptimisationSignal processingNumPy
mise
Shipped
Reads a clip's structure (composition, tone, rhythm, staging, motion) and finds public-domain films that share each axis.

The axes are measured near-independent, so matches are reported per axis, never as one style score. Descriptive, not a ranking.

VideoFilmRetrieval
02

READING ROOM

Clinical imaging is where calibration stops being academic.

I trained as a physician before I trained models. On a CT read, a confident wrong answer has a patient attached, so my imaging work is about what a model actually encodes and whether its confidence survives a new scanner, site or population.

  • 3D CT foundation models · DIRAC, parameter-matched against CNN and ViT on KiTS23, MSD and AbdomenAtlas
  • A locked CT pool · 6,212 abdomen and chest cases, QC'd, deduplicated, split with zero leakage
  • Label-free SSL · covtoken, coverage-constrained token pruning for medical images
  • Lung CT · probing frozen MedDINOv3 features for lesion identity on RIDER

Drag to window · ↔ width · ↕ level

03

RULES OF PLAY

  1. 01

    Write the success criteria before the data can argue back.

  2. 02

    A null result reported beats a positive one massaged.

  3. 03

    Match parameters, or it's a capacity result wearing a structure costume.

  4. 04

    Probe on CPU before spending GPU.

04

CONTINUE?

Building something that needs its confidence to hold up? Player 2 welcome.

Open to data science and ML roles · remote-friendly