DeepSeek releases V4 Flash Vision experimental model
DeepSeek published DeepSeek-V4-Flash-Vision-Exp on Hugging Face, an image-text-to-text model under the MIT license. The repository lists 866,024 downloads and 925 likes.
Linked and quoted from the labs, the press and the people using them. Times in UTC.
The 17 labs we track and the biggest stories beyond them. The rest is in Everything.
DeepSeek published DeepSeek-V4-Flash-Vision-Exp on Hugging Face, an image-text-to-text model under the MIT license. The repository lists 866,024 downloads and 925 likes.
Anthropic previewed the Model Hardware Standard, a standard that lets AI agents control machines such as microscopes. It was developed with medical research institute HHMI and is accessible to a limited number of users so far.
Google announced Gemini Omni 1.1 Flash, a new model in the Gemini Omni line. The release is described as offering more control for builders.
Tencent released Hy4-preview, a text-generation model, on Hugging Face. The listing tags it as a conversational mixture-of-experts model under the Hunyuan family.
Alibaba published the Qwen-Drive-1.0-4B model on Hugging Face. It is an image-text-to-text model for autonomous driving, motion planning, 3D perception and visual question answering.
Tencent released ContextPilot-8B, an 8B text-generation model built on Qwen3, on Hugging Face. The model targets context management, tool use, and agentic conversational tasks.
Cohere introduced Parse, a product for enterprise document intelligence at scale. The announcement was published on Cohere's official site.
Google released Gemini 3.5 Transcribe, a speech-to-text transcription model. The announcement says it provides more intelligent transcription.
IBM released its Granite 4.2 language models in 3B, 8B, and 30B sizes, trained on about 15 trillion tokens with a context window of up to 512,000 tokens. The larger models use agentic RL training to learn tool use and code execution, and all models are available under the Apache 2.0 license.
Zai released the GLM-5.3 model on Hugging Face. The model card lists text generation, conversational use, and support for English and Chinese.
Tencent published the WeMM-Embedding-2B model on Hugging Face. It is a multimodal embedding model for text and image embedding, built on qwen3_5.
Google released the TimesFM 3.0 PyTorch model for time-series forecasting on Hugging Face. The model has over 1.1 million downloads and 876 likes.
Alibaba published the Qwen3.8-Flash-Next image-text-to-text model on Hugging Face. The model card lists transformers and safetensors support, conversational use, eval results, and a custom license.
DeepSeek published API documentation for a vision model called DeepSeek-v4-flash-vision-exp. The model appears to add image understanding to the v4 flash line.
Google released the tipsv1-s14 model on Hugging Face. It is a vision model for zero-shot image classification and feature extraction.
Tencent's Hy-MT2-30B-A3B translation model is now listed on OpenRouter. It supports 33 language pairs plus five Chinese dialect and minority-language pairs.
OpenAI is shipping a version of ChatGPT tailored to users aged 13 to 17.
Tencent published the AuK model on Hugging Face for text-to-speech, voice cloning, speech generation, editing and enhancement. The model has 3895 downloads and 353 likes.
Tencent released the EVIE-Preview-4.5B model on Hugging Face. It is a vision-language model for visual document retrieval built on the colpali-engine.
Google Gemini and Pixel announced a partnership with five global football clubs. The partnership aims to improve the fan matchday experience using AI and smartphone technology.