YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows

pCoMole_overview_10s

Repository layout

pCoMole/
  train.py / evaluate.py          # Edit Flow train / val-loss eval
  configs/                        # GFP (protein) and SELFIES configs
  model/ logic/ flow_matching/    # Edit Flow implementation
  smiles_tokenizer/               # SMILES SPE + SELFIES vocab
  data/selfies/28k_mimetics/      # shipped SELFIES training split
  gfp/                            # GFP generation + pCoMole
    pcomole.py                    # official GFP editor
    generate.py
    ckpt/last_2.ckpt
    classifier_ckpt/best.pt
    FPredX/                       # excitation / brightness / emission models
  peptidomimetics/                # peptidomimetic generation + pCoMole
    pcomole.py
    generate.py
    objectives.py                 # PeptiVerse + Admetica/DeepDTAGen mix
    ckpt/SELFIES_EditFlows.ckpt
    ckpt/SMILES_BindEvaluator.ckpt
    ckpt/admetica/                # LD50, solubility, Caco-2, half-life
    ckpt/deepdtagen/
  cas9/                           # Cas9 generation + pCoMole
    pcomole.py                    # official Cas9 editor
    generate_cas9_csvs.py
    objectives.py                 # Cas9Classification, DeletionCount, PAMMatching
    constraints.py
    domain_detector.py            # HNH / RuvC completeness (HMMER)
    pam_detector_hmm.py           # PAM-interacting domain masking
    model/ logic/                 # Cas9 Edit Flow variant (see note below)
    data/orthologs.fasta          # 10 Cas9 ortholog seeds
    ckpt/last.ckpt
    classifier_ckpt/              # Cas9-likeness classifier
    hmm/                          # HMM databases

Note. cas9/model/ and cas9/logic/ are a fork of the repo-root packages, not duplicates: the Cas9 stack keeps the reparameterized Edit Flow models inside base_models.py, whereas the root splits them into reparam_models.py. The two are not interchangeable, so the Cas9 task imports its own copies (from cas9.model...).

1. Environment

A CUDA GPU is strongly recommended.

git clone <this-repo-url> pCoMole
cd pCoMole
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

HuggingFace model downloads happen on first use (facebook/esm2_t33_650M_UR50D, aaronfeller/PeptideCLM-23M-all, ChemBERTa).

PeptiVerse (required for peptidomimetic pCoMole)

Peptidomimetic property oracles use the latest PeptiVerse SMILES predictors. Clone it next to this repo, or set PEPTIVERSE_ROOT:

# recommended: sibling of pCoMole
cd ..
git clone https://huggingface.co/ChatterjeeLab/PeptiVerse
cd PeptiVerse
pip install -r requirements.txt
# follow PeptiVerse README to download model weights

pCoMole looks for PeptiVerse in this order:

  1. $PEPTIVERSE_ROOT
  2. ../PeptiVerse relative to this repository

PeptiVerse scores the peptide-like side of each candidate. Admetica + DeepDTAGen still score the small-molecule side, and the original peptide-likeness mix logic is unchanged.

MAFFT (required for GFP pCoMole)

GFP excitation, brightness, and emission oracles run through FPredX and need MAFFT on your PATH.

# conda
conda install -c bioconda mafft

# or point to a local binary
export MAFFT_PATH=/path/to/mafft

Confirm with mafft --version before running GFP editing.

HMMER and protein2pam (required for Cas9 pCoMole)

The Cas9 domain-completeness and PAM-masking oracles shell out to hmmscan:

conda install -c bioconda hmmer

The PAM objective needs the protein2pam model class (weights download from the Hub as Profluent-Bio/protein2pam-cas9_full):

pip install protein2pam==0.2.0

2. Train, evaluate, and sample Edit Flows

Run every command from the repository root.

SELFIES peptidomimetic Edit Flow

Training data is shipped at data/selfies/28k_mimetics.

# train
python train.py --config configs/config_selfies.yaml
# optional: python train.py --config configs/config_selfies.yaml --wandb

# evaluate a checkpoint on the validation split
python evaluate.py \
  --config configs/config_selfies.yaml \
  --ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt

# unconditional / seeded generation
python peptidomimetics/generate.py \
  --config configs/config_selfies.yaml \
  --ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt \
  --input 'CSCC[C@H](NC(=O)[C@@H]1CCCN1C(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@H](CCCNC(=N)N)NC(=O)[C@H](CO)NC(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H](N)CO)C(=O)N[C@@H](CC(=O)O)C(=O)O' \
  --num_steps 30 \
  --num_samples 2 \
  --output_csv outputs/peptidomimetic_unconditional.csv

Or: bash peptidomimetics/scripts/train.sh and bash peptidomimetics/scripts/generate.sh.

GFP protein Edit Flow

A trained GFP checkpoint is shipped at gfp/ckpt/last_2.ckpt. The original GFP training arrows are not included. To retrain, put a HuggingFace load_from_disk dataset under data/gfp/{train,validation} (or edit configs/config_gfp.yaml).

python evaluate.py --config configs/config_gfp.yaml --ckpt gfp/ckpt/last_2.ckpt

python gfp/generate.py \
  --config configs/config_gfp.yaml \
  --ckpt gfp/ckpt/last_2.ckpt \
  --input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS' \
  --num_steps 10 \
  --num_samples 8 \
  --output_csv outputs/gfp_unconditional.csv

3. pCoMole multi-objective editing

GFP

Official editor: gfp/pcomole.py (length + excitation + brightness, with GFP-classifier and emission constraints).

bash gfp/scripts/pcomole.sh

Equivalent command:

python gfp/pcomole.py \
  --root_dir gfp/FPredX \
  --config configs/config_gfp.yaml \
  --ckpt gfp/ckpt/last_2.ckpt \
  --num_steps 10 \
  --num_candidates 50 \
  --num_rollouts 10 \
  --objective_weights 3 1 1 \
  --output_file outputs/gfp_length_excitation_brightness.csv \
  --input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS'

If the output filename contains length_excitation_brightness, length_brightness, length_excitation, or length, the matching objective subset is used. Any other name uses all three objectives.

Peptidomimetics

Official editor: peptidomimetics/pcomole.py.

Objectives (7 scores when --specificity is on):

  1. non-toxicity
  2. solubility
  3. permeability
  4. half-life
  5. affinity
  6. motif
  7. specificity

PeptiVerse provides the peptide-side scores. Admetica + DeepDTAGen provide the small-molecule side. The two are mixed by peptide-likeness, same as the paper code. Motif / specificity still use the shipped BindEvaluator checkpoint.

# after PeptiVerse is cloned and its weights are downloaded
bash peptidomimetics/scripts/pcomole.sh

--target and --motifs are required for affinity and motif oracles. --specificity adds the seventh objective.

Cas9

Official editor: cas9/pcomole.py. Three objectives โ€” Cas9-likeness, deletion count, and PAM matching โ€” under Cas9-score, PAM-matching and length constraints.

Cas9 code is a package, so run it from the repository root with -m:

ORTHOLOG=cjcas9 bash cas9/scripts/pcomole.sh

Equivalent command:

python -m cas9.pcomole \
  --pcomol_config configs/config_cas9.yaml \
  --ckpt cas9/ckpt/last.ckpt \
  --input "$(awk -v n='>cjcas9' '$1==n{f=1;next}/^>/{f=0}f{printf "%s",$0}' cas9/data/orthologs.fasta)" \
  --cas9_objective --deletion_count_obj --pam_distr_objective \
  --objective_weights 1 1 1 \
  --target_pam matching \
  --num_steps 5 --num_rollouts 5 --num_candidates 10 --num_sequences 50 \
  --max_deletion_percentage 0.2 --deletion_rate_scale 2500 \
  --pam_min_confidence 0.7 --cas9_score_threshold 0.5 \
  --max_target_length 930 --min_target_length 890 \
  --use_ce_loss --allow_equal_length --zero_lam_ins \
  --output_file outputs/cas9_cjcas9_pcomole.csv

Ten ortholog seeds ship in cas9/data/orthologs.fasta (SpCas9, SaCas9, CjCas9, St1/St3Cas9, GeoCas9, NmCas9, iSpyMac, SpCas9++, SpRYc). Set ORTHOLOG to any of them. --target_pam matching optimizes toward the seed's own PAM; pass an explicit 10-nt ACGTN string to retarget.

4. Included checkpoints

Path Role
gfp/ckpt/last_2.ckpt GFP Edit Flow
gfp/classifier_ckpt/best.pt GFP hard constraint
gfp/FPredX/{ex,bright,em}_model FPredX property models
peptidomimetics/ckpt/SELFIES_EditFlows.ckpt peptidomimetic Edit Flow
peptidomimetics/ckpt/SMILES_BindEvaluator.ckpt motif / specificity
peptidomimetics/ckpt/admetica/*.ckpt Admetica blending oracles
peptidomimetics/ckpt/deepdtagen/ DeepDTAGen affinity
cas9/ckpt/last.ckpt Cas9 Edit Flow
cas9/classifier_ckpt/ Cas9-likeness classifier (objective + hard constraint)
cas9/hmm/ HNH / RuvC and PAM-interacting HMM databases

Together these are about 10 GB. HuggingFace uploads should use Git LFS for the .ckpt / .pt / .pth files.

5. Environment variables

Variable Purpose
PEPTIVERSE_ROOT PeptiVerse repo root (manifest + training_classifiers/)
MAFFT_PATH MAFFT binary, if it is not on PATH
GFP_CLASSIFIER_CKPT optional override for the GFP classifier
PROTEIN2PAM_ROOT protein2pam package root (PAM oracle model class)
CAS_PREDICTOR_ROOT cas_predictor package root (Cas9 classifier code)
CAS9_CLASSIFIER_CKPT override for the Cas9 classifier checkpoint
CAS9_CLASSIFIER_CONFIG optional Cas9 classifier config
CAS9_HMM_DIR directory holding the Cas9 HMM databases

6. Notes

  • Run scripts from the repo root, or use the wrappers under */scripts/.
  • GFP FPredX will fail immediately if mafft is missing; that is expected.
  • Peptidomimetic pCoMole will fail at oracle init if PeptiVerse is missing or its weights were not downloaded.
  • Cas9 objectives fail at init if hmmscan, protein2pam, or cas_predictor are missing.
  • Training writes to outputs/ by default.

Citation

If you use this code, please cite the pCoMole paper and PeptiVerse.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support