#多模态
Posts connected by this tag.
2 postsView all tags →
From DeepSeek visual primitives and MiniMax-M3 image TTFT to how MiniCPM-o sustains local video understanding and full-duplex speech interaction.
3346 words
|
17 minutes

An investigation into why web and API results diverge during multi-image analysis, from Lost in the Middle to a two-stage approach based on per-image pre-summaries.
2061 words
|
10 minutes

