Multimodal models
Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.
Fine-Tune MMS Adapter Models for low-resource ASR
AI Speech Recognition in Unity
GPT-4
We’ve created GPT-4, the latest milestone in OpenAI’s effort in scaling up deep learning. GPT-4 is a large multimodal model (accepting image and text inputs, emitting text outputs) that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks.
A Dive into Vision-Language Models
Fine-Tune Whisper For Multilingual ASR with 🤗 Transformers
Introducing Whisper
Making automatic speech recognition work on large files with Wav2Vec2 in 🤗 Transformers
Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with 🤗 Transformers
Fine-Tune Wav2Vec2 for English ASR in Hugging Face with 🤗 Transformers
Multimodal neurons in artificial neural networks
We’ve discovered neurons in CLIP that respond to the same concept whether presented literally, symbolically, or conceptually. This may explain CLIP’s accuracy in classifying surprising visual renditions of concepts, and is also an important step toward understanding the associations and biases that CLIP and similar models learn.