Models
Datasets
Spaces
Docs
Enterprise
Pricing
Log In
Sign Up

Collections

Discover the best community collections!

Collections including paper arxiv:2308.13418

Language agents achieve superhuman synthesis of scientific knowledge

Paper • 2409.13740 • Published Sep 10, 2024
Nougat: Neural Optical Understanding for Academic Documents

Paper • 2308.13418 • Published Aug 25, 2023 • 41
datalab-to/chandra

Image-to-Text • 9B • Updated Oct 21 • 87.1k • 409
Running

25

Denario

😻

25

GUI for Denario

LayerCake: Token-Aware Contrastive Decoding within Large Language Model Layers

Paper • 2507.04404 • Published Jul 6 • 21
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float

Paper • 2504.11651 • Published Apr 15 • 31
A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone

Paper • 2505.12781 • Published May 19 • 2
A Survey of Context Engineering for Large Language Models

Paper • 2507.13334 • Published Jul 17 • 259

LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Paper • 2306.17107 • Published Jun 29, 2023 • 11
On the Hidden Mystery of OCR in Large Multimodal Models

Paper • 2305.07895 • Published May 13, 2023 • 1
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities

Paper • 2308.12966 • Published Aug 24, 2023 • 11
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Paper • 2401.15947 • Published Jan 29, 2024 • 53

Nougat: Neural Optical Understanding for Academic Documents

Paper • 2308.13418 • Published Aug 25, 2023 • 41

Language models are weak learners

Paper • 2306.14101 • Published Jun 25, 2023 • 10
Large Language Models as Tax Attorneys: A Case Study in Legal Capabilities Emergence

Paper • 2306.07075 • Published Jun 12, 2023 • 10
TableGPT: Towards Unifying Tables, Nature Language and Commands into One GPT

Paper • 2307.08674 • Published Jul 17, 2023 • 48
Nougat: Neural Optical Understanding for Academic Documents

Paper • 2308.13418 • Published Aug 25, 2023 • 41

Awesome Document AI

A collection of open-source document AI 📄 📝 📈

Runtime error

Featured

84

UDOP

🏃

84

Generate text from document images
Runtime error

40

Pix2struct

📚

40

Play with all the pix2struct variants in this d
Sleeping

26

Compare Docvqa Models

🦀

26

Compare different visual question answering
Runtime error

Featured

289

DocQuery — Document Query Engine

🦉

289

Nougat: Neural Optical Understanding for Academic Documents

Paper • 2308.13418 • Published Aug 25, 2023 • 41
Kosmos-2.5: A Multimodal Literate Model

Paper • 2309.11419 • Published Sep 20, 2023 • 55

Nougat: Neural Optical Understanding for Academic Documents

Paper • 2308.13418 • Published Aug 25, 2023 • 41
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Paper • 2307.02499 • Published Jul 4, 2023 • 15
Text Rendering Strategies for Pixel Language Models

Paper • 2311.00522 • Published Nov 1, 2023 • 12

Language agents achieve superhuman synthesis of scientific knowledge

Paper • 2409.13740 • Published Sep 10, 2024
Nougat: Neural Optical Understanding for Academic Documents

Paper • 2308.13418 • Published Aug 25, 2023 • 41
datalab-to/chandra

Image-to-Text • 9B • Updated Oct 21 • 87.1k • 409
Running

25

Denario

😻

25

GUI for Denario

Language models are weak learners

Paper • 2306.14101 • Published Jun 25, 2023 • 10
Large Language Models as Tax Attorneys: A Case Study in Legal Capabilities Emergence

Paper • 2306.07075 • Published Jun 12, 2023 • 10
TableGPT: Towards Unifying Tables, Nature Language and Commands into One GPT

Paper • 2307.08674 • Published Jul 17, 2023 • 48
Nougat: Neural Optical Understanding for Academic Documents

Paper • 2308.13418 • Published Aug 25, 2023 • 41

LayerCake: Token-Aware Contrastive Decoding within Large Language Model Layers

Paper • 2507.04404 • Published Jul 6 • 21
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float

Paper • 2504.11651 • Published Apr 15 • 31
A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone

Paper • 2505.12781 • Published May 19 • 2
A Survey of Context Engineering for Large Language Models

Paper • 2507.13334 • Published Jul 17 • 259

Awesome Document AI

A collection of open-source document AI 📄 📝 📈

Runtime error

Featured

84

UDOP

🏃

84

Generate text from document images
Runtime error

40

Pix2struct

📚

40

Play with all the pix2struct variants in this d
Sleeping

26

Compare Docvqa Models

🦀

26

Compare different visual question answering
Runtime error

Featured

289

DocQuery — Document Query Engine

🦉

289

LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Paper • 2306.17107 • Published Jun 29, 2023 • 11
On the Hidden Mystery of OCR in Large Multimodal Models

Paper • 2305.07895 • Published May 13, 2023 • 1
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities

Paper • 2308.12966 • Published Aug 24, 2023 • 11
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Paper • 2401.15947 • Published Jan 29, 2024 • 53

Nougat: Neural Optical Understanding for Academic Documents

Paper • 2308.13418 • Published Aug 25, 2023 • 41
Kosmos-2.5: A Multimodal Literate Model

Paper • 2309.11419 • Published Sep 20, 2023 • 55

Nougat: Neural Optical Understanding for Academic Documents

Paper • 2308.13418 • Published Aug 25, 2023 • 41

Nougat: Neural Optical Understanding for Academic Documents

Paper • 2308.13418 • Published Aug 25, 2023 • 41
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Paper • 2307.02499 • Published Jul 4, 2023 • 15
Text Rendering Strategies for Pixel Language Models

Paper • 2311.00522 • Published Nov 1, 2023 • 12

Company

TOS Privacy About Careers

Website

Models Datasets Spaces Pricing Docs