Projects based on SigLIP (Zhai et. al, 2023) and Hugging Face transformers integration 🤗
-
Updated
Feb 21, 2025 - Jupyter Notebook
Projects based on SigLIP (Zhai et. al, 2023) and Hugging Face transformers integration 🤗
Inference and fine-tuning examples for vision models from 🤗 Transformers
[ICCVW 25] LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
[CVPR'25-Demo] Official repository of "TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models".
[NeurIPS 2024] AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
Local-first cinematic visual archive for filmmakers — search stills by look, mood & technique; ingest video/URLs; moodboards. Own your frames (FilmGrab / Flim / Kive style, on your machine).
本项目以应用为主出发,结合了从基础的机器学习、深度学习到目标检测以及目前最新的大模型,采用目前成熟的 第三方库、开源预训练模型以及相关论文的最新技术,目的是记录学习的过程同时也进行分享以供更多人可以直接进行使用。
[ICLR 2026] The implementation of the paper Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
[ICLR 2025] - Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Fine-Tuning SigLIP 2 for Single/Multi-Label Image Classification. Image classification vision-language encoder model fine-tuned for Image Classification Tasks
Low-latency ONNX and TensorRT based zero-shot classification and detection with contrastive language-image pre-training based prompts
Download flickr8k, flickr30k image caption datasets
Official PyTorch implementation of the WACV 2025 Oral paper "Composed Image Retrieval for Training-FREE DOMain Conversion".
A minimal, but effective implementation of CLIP (Contrastive Language-Image Pretraining) in PyTorch
Code for Post-hoc Probabilistic Vision-Language Models
Chitrarth: Bridging Vision and Language for a Billion People
Meme search and discovery engine using OpenAI CLIP and Salesforce BLIP
Open reproducible benchmarks for food-image recognition models and APIs.
CLIP & SigLIP model training from scratch
An AI powered Video Serach Engine with google's SigLIP and Qdrant. It allows to search objects or key moments in videos just using natural language.
Add a description, image, and links to the siglip topic page so that developers can more easily learn about it.
To associate your repository with the siglip topic, visit your repo's landing page and select "manage topics."