Phando
/

chemberta-v2-finetuned-uspto-50k-classification

Text Classification

Inference Endpoints

Model card Files Files and versions Community

This ChemBERTa-v2 checkpoint was fine-tuned on the USPTO-50k dataset for sequence classification.

Specifically, the objective is to predict the reaction class label, and the input is either (canonicalized) all reactant SMILES or all product SMILES (separated by ".").

Train/Test split: 0.99/0.01
Evaluation results:
- Accuracy: 87.11%
- Loss: 0.4272
Fine-tuning hyperparameters:
- seed = 233
- batch-size = 128
- num_epochs = 5 (but early stopped at epoch 4)
- learning_rate = 5e-4
- warmup_steps = 64
- weight_decay = 0.01
- lr_scheduler_type = "cosine"

Downloads last month: 79

Safetensors

Model size

83.5M params

Tensor type

F32

·

Inference Examples

Text Classification

This model does not have enough activity to be deployed to Inference API (serverless) yet. Increase its social visibility and check back later, or deploy to Inference Endpoints (dedicated) instead.

Dataset used to train Phando/chemberta-v2-finetuned-uspto-50k-classification