AI & ML interests
Open models built to place at the top of their class on public leaderboards, from fine-tunes the open-source community has already published. Every result is measured against the base model and published with its evidence.
Recent Activity
Optitransfer
Open models built to place at the top of their class on public leaderboards, from the work the open-source community has already published.
At a Glance
The Thesis
Thousands of fine-tuned models sit on public repositories. Each adds a skill to a shared base model. Their training is already paid for.
Hypothesis: that capability can be consolidated into one model per class without pre-training, with trade-offs between skills controlled and capability then maximised across the board through iteration. Converge is the programme that tests it, one release at a time, in public.
Where We Are
- v2 against its base. 12 benchmarks: 2 improved, 2 regressed, 8 with no measurable difference.
- Improved. GSM8K under HELM's protocol: 87.0 against 83.3 (+3.7, calibrated). MMLU under HELM's protocol: +0.47 (calibrated).
- Regressed. MATH Level 5: -5.74. MMLU-Pro: -1.89.
- Defect found. In the merge that produced v1, inherited by v2. Corrected. The corrected v1 measures at its base model's level on MMLU-Pro (TIGER-Lab).
- Not yet shown. The end state: a release that improves on its base on every axis. Each iteration is built to recover the regressions and raise the rest.
Why It Matters
- Science. Whether trade-offs can be controlled, and then removed through iteration, is an open, testable question.
- Economics. If the hypothesis holds, the marginal cost of a stronger open model is evaluation compute, not pre-training.
- Adoption. Capability without dependence on one vendor, with a record of every model change.
The Goal
Top placements in class, across the public leaderboards.
- Scope. 7B first. Then each larger size class and other model families.
- Method. Each release controls the trade-offs between skills. Iteration then raises every axis: reasoning, mathematics, code, instruction following and knowledge.
- End state. Improvement on every axis, with no statistically significant regression against the base model.
- Done. A class is complete when no public fine-tune improves any axis further.
The Process
- Start from a strong open base model in the class
- Discover the compatible fine-tunes the community has published for it
- Assimilate them into one model
- Measure every axis against the base model, on identical items and public protocols
- Release with every gain and every regression published, against the base model and the previous release
- Iterate to recover the regressions and raise the weaker axes, until nothing more improves. Then the next class
The construction method is proprietary. The evaluation is public.
Releases
Published under this organisation from converge v3.
Research
The research behind the releases is on the founder's page, @Optitransfer.
The corrected models' boards are published when complete, whichever way they fall.
Evidence Standard
- Paired. Every score is the model minus its base on the same items, with a 95% confidence interval and an exact significance test
- Regressions reported. A card lists what went down as prominently as what went up
- Calibration labelled. Calibrated means the harness first reproduced the base model's published score. Every result is marked either way
- Public protocols. HELM-protocol GSM8K and MMLU, MMLU-Pro (TIGER-Lab), ZeroEval, EvalPlus HumanEval+ and MBPP+, MATH, AIME, ARC, IFEval
- Corrected in public. When a defect is found, the affected cards say so first
Open Questions
- Can trade-offs be controlled, and every axis then raised through iteration, across a full board?
- Does it beat ensembling, routing and best-of-n sampling at matched compute?
- Does it carry to larger classes and other model families?
- Does it hold on calibration, hallucination and instruction following?
Each answer is published as it is measured.
Who It Is For
- Open community. A stronger open model per class, with the evidence to check it
- Regulated industries. Healthcare and finance teams whose fine-tunes cannot leave their boundary, and who must show how each model change was made
- National programmes. Capability that grows by contribution, without dependence on a few vendors
Open model, paid guarantees. The public model stays open. The planned commercial offering is private consolidation of an organisation's own fine-tunes, inside its boundary, with the same measurement and evidence.