Change8
Symbol10 releases

AutoProcessor

Found in 1 package: transformers

transformers(10 releases)

4.55.0-GLM-4.5V-preview
Aug 11, 2025

This release introduces GLM-4.5V, a high-performance multimodal reasoning model based on GLM-4.5-Air, featuring advanced capabilities in image, video, and GUI analysis.

v4.51.3-CSM-previewBreaking
May 8, 2025

This release introduces the Conversational Speech Model (CSM), an open-source contextual text-to-speech model capable of generating natural speech from multi-turn dialogue context.

v4.51.3-LlamaGuard-preview
Apr 30, 2025

This release introduces LlamaGuard 4 and Llama Prompt Guard 2, providing multimodal safety moderation for text and images. It is available as a preview tag prior to the official v4.52.0 minor release.

v4.51.3-InternVL-preview
Apr 22, 2025

This preview release introduces support for the InternVL 2.5 and 3 family of multimodal models, featuring a native multimodal pre-training paradigm and state-of-the-art performance on visual-linguistic tasks.

v4.51.3-MLCD-preview
Apr 22, 2025

This release introduces a preview of the MLCD vision model, a foundational visual model optimized for multimodal LLMs like LLaVA, developed by DeepGlint-AI.

v4.51.0Breaking
Apr 5, 2025

This release introduces support for Llama 4, Phi4-Multimodal, DeepSeek-v3, and Qwen3 architectures, alongside a major documentation overhaul and modularization of speech models.

v4.49.0-Mistral-3
Mar 18, 2025

This release introduces Mistral 3 (Mistral Small 3.1) to the Transformers library, a 24B parameter model featuring 128k context length and advanced vision-language capabilities.

v4.49.0-Gemma-3
Mar 18, 2025

This release introduces Google's Gemma 3 multimodal models to the transformers library, featuring a SigLIP vision encoder and Gemma 2 language decoder with support for high-resolution image cropping and multi-image inference.

v4.49.0-AyaVision
Mar 4, 2025

This release introduces Aya Vision 8B and 32B, multilingual multimodal models combining SigLIP-2 vision encoders with Cohere language models, available via a specialized transformers release tag.

v4.49.0-SmolVLM-2
Feb 20, 2025

This release introduces SmolVLM-2, a lightweight vision-language model based on Idefics3 and SmolLM2 that supports multi-image and video processing.

Track Symbol Changes

Use the Change8 MCP server or GitHub Action to get notified when AutoProcessor changes.

Learn More