#多模态

Posts connected by this tag.

From Single-Frame Vision to Real-Time Video Conversation: The Multimodal Architecture Paths of DeepSeek, MiniCPM, and MiniMax
From DeepSeek visual primitives and MiniMax-M3 image TTFT to how MiniCPM-o sustains local video understanding and full-duplex speech interaction.
3346 words
|
17 minutes
Cover Image of the Post
Single Image Does Not Equal Multiple Images: Why VLMs Hallucinate More with Multiple Images, and a Two-Stage Fix
An investigation into why web and API results diverge during multi-image analysis, from Lost in the Middle to a two-stage approach based on per-image pre-summaries.
2061 words
|
10 minutes
Cover Image of the Post