-
microsoft/trocr-large-handwritten
Image-to-Text • Updated • 57.1k • 167 -
google/gemma-4-31B-it
Image-Text-to-Text • 31B • Updated • 9.04M • • 3.87k -
TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models
Paper • 2109.10282 • Published • 13 -
HTR-VT: Handwritten Text Recognition with Vision Transformer
Paper • 2409.08573 • Published
Collections
Discover the best community collections!
Collections including paper arxiv:2109.10282
-
FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction
Paper • 2305.02549 • Published • 7 -
FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction
Paper • 2203.08411 • Published • 2 -
More efficient manual review of automatically transcribed tabular data
Paper • 2306.16126 • Published • 1 -
CascadeTabNet: An approach for end to end table detection and structure recognition from image-based documents
Paper • 2004.12629 • Published • 3
-
Can large language models explore in-context?
Paper • 2403.15371 • Published • 31 -
GaussianCube: Structuring Gaussian Splatting using Optimal Transport for 3D Generative Modeling
Paper • 2403.19655 • Published • 19 -
WavLLM: Towards Robust and Adaptive Speech Large Language Model
Paper • 2404.00656 • Published • 11 -
Enabling Memory Safety of C Programs using LLMs
Paper • 2404.01096 • Published • 1
-
microsoft/trocr-large-handwritten
Image-to-Text • Updated • 57.1k • 167 -
google/gemma-4-31B-it
Image-Text-to-Text • 31B • Updated • 9.04M • • 3.87k -
TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models
Paper • 2109.10282 • Published • 13 -
HTR-VT: Handwritten Text Recognition with Vision Transformer
Paper • 2409.08573 • Published
-
FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction
Paper • 2305.02549 • Published • 7 -
FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction
Paper • 2203.08411 • Published • 2 -
More efficient manual review of automatically transcribed tabular data
Paper • 2306.16126 • Published • 1 -
CascadeTabNet: An approach for end to end table detection and structure recognition from image-based documents
Paper • 2004.12629 • Published • 3
-
Can large language models explore in-context?
Paper • 2403.15371 • Published • 31 -
GaussianCube: Structuring Gaussian Splatting using Optimal Transport for 3D Generative Modeling
Paper • 2403.19655 • Published • 19 -
WavLLM: Towards Robust and Adaptive Speech Large Language Model
Paper • 2404.00656 • Published • 11 -
Enabling Memory Safety of C Programs using LLMs
Paper • 2404.01096 • Published • 1