大佬们,针对目前图片、PDF解析,含内容、内容中的图片、公式、表格等。有哪些效果好的开源模型。小弟在此蹲守学习…
现在 OCR 方面gemini遥遥领先
开源啊,那其实没几个选择了,ds系吧
Qwen VL 或 InternVL 或看看佬友发的帖子
OCR选百度的paddleocr就行
DS 不识图的
哦哦,可参考这个榜单
学习下,感谢~
榜单就是给人看着玩的 ![]()
刚发现一个MonkeyOcr,感觉好像还可以,准备部署下试试
PaddleOCR‑VL
是的,所以说参考
那不如试试 OCRFlux
最近有人开源了一个这个,可以试试
mineru
PaddleOCR‑VL 前两天发布,对于你的需求完美契合。
nice
补充
- GitHub - PaddlePaddle/PaddleOCR: Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
- GitHub - QwenLM/Qwen3-VL: Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
- GitHub - opendatalab/MinerU: Transforms complex documents like PDFs into LLM-ready markdown/JSON for your Agentic workflows.
- GitHub - rednote-hilab/dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
OCR选百度家的就行
