Explosion
Explosion builds developer tools for AI, Machine Learning and Natural Language Processing.
Project
Topics
Category
Tasks
Select...Code Generation, Coreference Resolution, Dependency Parsing, Distillation, Embeddings & Vectors, Entity Linking, Evaluation, Image Classification, Image Segmentation, Layout Analysis, Lemmatization, Named Entity Recognition, Object Detection, Optical Character Recognition (OCR), Part-of-Speech Tagging, PII Anonymization, Question Answering, Relation Extraction, Retrieval-Augmented Generation (RAG), Rule-Based Matching, Span Categorization, Text Classification, Text Generation, Tokenization, Weak Supervision.
Authors
Select...Adriane Boyd, Ákos Kádár, Basile Dura, Chung-Fan Tsai, Damian Romero, Daniel de Kok, Duygu Altinok, Edward Schmuhl, Helena Steckmeister, India Kerle, Ines Montani, Kabir Khan, Lj Miranda, Madeesh Kannan, Magdalena Anioł, Matthew Honnibal, Paul O’Leary McCann, Peter Baumgartner, Philip Vollet, Raphael Mitsch, Rehan Ahmed, Richard Hudson, Ryan Wesslen, Sofie Van Landeghem, Victoria Slocum, Vincent D. Warmerdam, Vinit Ravishankar, Walter Henry.
Conquering PDFs: document understanding beyond plain text
In this talk, Ines presents a new and modular approach for building robust document understanding systems, using state-of-the-art models and the awesome Python ecosystem.
Microsoft Presidio v2.2.352
Context aware, pluggable and customizable PII de-identification and anonymization service for text and images, featuring a spaCy back-end.
Finding Bad Image Data using UMAP and Prodigy
In this video, we’ll show you how to use Prodigy to find bad examples in the Google QuickDraw dataset. We will be leveraging a technique that involves UMAP to find strange images semi-automatically.
Prodigy-Segment for Pixel Segmentation
Use Meta’s “Segment Anything” model in Prodigy to help you select the right pixels in images.
Prodigy v1.10: Dependencies, relations, audio, video & more
Version 1.10 of Prodigy includes tons of new features, including manual dependency and relation annotation, audio and video annotation, a new and improved image UI, new recipe callbacks, more settings for manual NER, plus various new config options and settings.
Best Way to OCR a PDF in Python
Tutorial by WJB Mattingly on how to use the new spaCy Layout package and Docling to convert PDFs to text.
Image Captioning with Prodigy & PyTorch
In this video, we’ll show you how you can use Prodigy to script fully custom annotation workflows in Python, how to plug in your own machine learning models and how to mix and match different interfaces for your specific use case.
From PDFs to AI-ready structured data: a deep dive
This blog post presents a new modular workflow for converting PDFs and similar documents to structured data and shows you how to build end-to-end document understanding and information extraction pipelines for industry use cases.
Prodigy-PDF for PDF annotation and OCR
Want to annotate PDF files? Our new Prodigy plugin can help with that! To explain how to use PDF segmentation and OCR, Vincent made a small demo video.
Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
We use the spaCy library for tokenization, part-of-speech tagging, and lemmatization of the words in the descriptions.
Finetuning and Bulk Labelling Images with Prodigy
In this video, we’ll show how you might be able to improve the annotation experience by using bulk labelling for image classification.