Efficient Track Anything
[🤗Checkpoints
][📕Project
][🤗Gradio Demo
][📕Paper
]
The Efficient Track Anything Model(EfficientTAM) takes a vanilla lightweight ViT image encoder. An efficient memory cross-attention is proposed to further improve the efficiency. Our EfficientTAMs are trained on SA-1B (image) and SA-V (video) datasets. EfficientTAM achieves comparable performance with SAM 2 with improved efficiency. Our EfficientTAM can run >10 frames per second with reasonable video segmentation performance on iPhone 15. Try our demo with a family of EfficientTAMs at [🤗Gradio Demo
].
This repository contains a family of EfficientTAMs with checkpoints for practical deployments with different latency and quality needs.