Vela commitment classifier — experimental pilot

A short English statement classifier that detects an expressed decision, firm intention, or commitment. This is a personal fine-tuning experiment derived from Vela, not an official vLLM Semantic Router release or an endorsed product.

Label Meaning
DECISION_EXPRESSED An affirmed chosen action, firm intention, commitment, explicit rejection or chosen inaction. Explicitly reported past choices also count.
NO_DECISION_EXPRESSED Preferences, questions, requests, possibilities, predictions, example quotations, or completed-action reports without an expressed choice.

It does not determine whether a promise was fulfilled, infer an unspoken mental state, extract tasks, or decide whether a statement deserves a decision-log entry. Routine firm process commitments count as positive under this broad contract.

Quick start

pip install torch transformers==4.57.6
from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="TwilightTechie/vela-commitment-classifier-307m",
    tokenizer="TwilightTechie/vela-commitment-classifier-307m",
    device=-1,  # CPU
)
print(classifier("We have decided to update the document today."))
print(classifier("I like mountains."))

The first example was classified as DECISION_EXPRESSED; the second as NO_DECISION_EXPRESSED. A GPU or hosted endpoint is not required for inference. Scores are model softmax outputs, not calibrated certainty. Run only trusted standard library code; this model requires no trust_remote_code=True.

Keep inputs to 256 tokens including special tokens. You can enforce this before classification:

text = "We agreed to postpone the rollout until the pilot finishes."
if len(classifier.tokenizer(text)["input_ids"]) > 256:
    raise ValueError("This pilot's input budget is 256 tokens")
print(classifier(text))

The foundation has a larger architectural context capacity; this fine-tuned pilot has not established quality on long documents or conversations.

Training and provenance

  • Foundation: vllm-sr/Vela-1.0-Encoder-307M.
  • Exact foundation revision: 720ab37904ce15054068d429381bfae7549f00f6.
  • Architecture: ModernBERT sequence classifier with a fresh two-label task head.
  • Method: complete checkpoint, updating the final two encoder layers and task head; other encoder parameters stayed frozen and passed preservation checks.
  • Trainable parameters: 10,622,210.
  • Training data: 118 short, synthetic English examples. Development data: 14. Assistant-authored examples were accepted by the user for this pilot; this does not constitute independent expert annotation. Original row review and provenance fields remain unchanged in the dataset.
  • 100 optimizer steps, seed 42, batch size 2, gradient accumulation 2.
  • Encoder learning rate 0.00002; head learning rate 0.0001; maximum length 256.
  • BF16 forward autocast, FP32 parameters; FP32 development evaluation.
  • Hardware: one NVIDIA L4 on Modal; peak allocated training memory 1.3697 GiB.
  • Selection: development macro F1 at steps 0, 25, 50, 75, 100. Step 75 selected.
  • Remote environment: Python 3.12, Torch 2.8.0+cu128, Transformers 4.57.6, PEFT 0.18.1, Accelerate 1.10.1, scikit-learn 1.7.2.

The training recipe and environment are included under reproducibility/. The training code comes from the public vLLM Semantic Router repository. source-files.json captures hashes of the exact trainer modules used; the experiment was made in a local checkout, not claimed as an upstream release. No private source document or credentials are included.

Development results — not independent testing

All rows below use the same 14-example development set:

Classifier Correct Macro F1
Fixed keyword rule 12/14 0.8542
TF-IDF + logistic regression 13/14 0.9282
Decision 2.0 Kai, fixed yes/no question 11/14 0.7754
This selected Vela checkpoint 14/14 1.0000

The neural checkpoint was selected using this development set. Additional training examples were authored after inspecting earlier development errors. Therefore, 14/14 does not imply general accuracy or enterprise readiness. FP32 GPU and local CPU reload produced the same class predictions; maximum probability difference was about 0.00000185. This is artifact verification, not independent quality evaluation.

A 16-example same-author synthetic test set remains unpublished and unevaluated. A fresh, realistic, independently annotated evaluation set is still needed.

Intended use and limitations

Use this for learning, local experimentation, and a candidate classifier in human-reviewed workflows. It has not been qualified for automatic authoritative logging, production routing, or consequential decisions.

  • English-only evaluated scope despite the multilingual foundation.
  • Very small authored dataset; wording and topic bias are likely.
  • Standalone short statements only; multi-turn reference resolution is not tested.
  • Conditional, canceled, sarcastic and context-dependent statements need a separately reviewed contract; predictions may be unreliable there.
  • A past decision may be superseded; this classifier does not identify current policy, author authority, or supersession links.
  • High scores can be overconfident. Calibration and realistic false-positive rates have not been measured.
  • Transformers 4.57.6 emitted a known class of Mistral regex warnings when loading this non-Mistral tokenizer locally. No tokenizer patch was applied. Basic reload checks passed; a broader tokenizer audit has not been performed.

License and attribution

The pinned Vela model card declares MIT; it identifies jhu-clsp/mmBERT-base as its upstream foundation, whose model card also declares MIT. The fine-tuned model is shared under MIT with those upstream declarations retained. See LICENSE and UPSTREAM_LICENSES.md for the recorded absence of separate upstream notice files. No upstream copyright owner or year has been invented.

The reproducibility code and synthetic dataset are distributed separately under Apache-2.0, following the Semantic Router repository's code/data context. See reproducibility/LICENSE and the dataset card. This model license does not relicense upstream training corpora or unrelated dependencies.

Downloads last month
3
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TwilightTechie/vela-commitment-classifier-307m

Finetuned
(14)
this model

Dataset used to train TwilightTechie/vela-commitment-classifier-307m