Explosion builds developer tools for AI, Machine Learning and Natural Language Processing.

Project

Topics

Tasks

Select...Code Generation, Coreference Resolution, Dependency Parsing, Distillation, Embeddings & Vectors, Entity Linking, Evaluation, Image Classification, Image Segmentation, Layout Analysis, Lemmatization, Named Entity Recognition, Object Detection, Optical Character Recognition (OCR), Part-of-Speech Tagging, PII Anonymization, Question Answering, Relation Extraction, Retrieval-Augmented Generation (RAG), Rule-Based Matching, Span Categorization, Text Classification, Text Generation, Tokenization, Weak Supervision

Authors

Select...Adriane Boyd, Ákos Kádár, Basile Dura, Chung-Fan Tsai, Damian Romero, Daniel de Kok, Duygu Altinok, Edward Schmuhl, Helena Steckmeister, India Kerle, Ines Montani, Kabir Khan, Lj Miranda, Madeesh Kannan, Magdalena Anioł, Matthew Honnibal, Paul O’Leary McCann, Peter Baumgartner, Philip Vollet, Raphael Mitsch, Rehan Ahmed, Richard Hudson, Ryan Wesslen, Sofie Van Landeghem, Victoria Slocum, Vincent D. Warmerdam, Vinit Ravishankar, Walter Henry

Beta test our upcoming product for agentic NLP

We’re looking for beta partners for Ellf: a platform and virtual assistant that makes your coding agent like Claude Code proficient at developing NLP solutions. Contact us or sign up for the waitlist if your team needs help with a project focused on tasks like information extraction!

How Love Without Sound helps the music industry recover millions in revenue for artists with NLP, spaCy and Prodigy

A case study on Love Without Sound’s innovative AI-powered tools for the music industry and law firms specializing in royalty negotiations.

How GitLab uses spaCy to analyze support tickets and empower their community

A case study on GitLab’s large-scale NLP pipelines for extracting actionable insights from support tickets and usage questions.

How S&P Global is making markets more transparent with NLP, spaCy and Prodigy

A case study on S&P Global’s efficient information extraction pipelines for real-time commodities trading insights in a high-security environment.

Introducing spaCy v3.6

spaCy v3.6 introduces the span finder component and trained pipelines for Slovenian.

The Tale of Bloom Embeddings and Unseen Entities

The default Bloom embedding layer in spaCy is unconventional, but very powerful and efficient. We wrote about it before and showed the advantages it provides in terms of memory efficiency for our floret embeddings. Now we have released the first technical report by Explosion, where we explain Bloom embeddings in more detail and rigorously compare them to traditional embeddings.

Explosion in 2022: Our Year in Review

It's been another exciting year at Explosion! We've developed a new end-to-end neural coref component for spaCy, improved the speed of our CNN pipelines up to 60%, and published new pre-trained pipelines for Finnish, Korean, Swedish and Croatian. We've also released several updates to Prodigy and introduced new recipes to kickstart annotation with zero- or few-shot learning.

spaCy Cheat Sheet

Everything you need to know about spaCy as a handy two-page PDF.

Introducing Holmes 4.0

A few weeks ago we released version 4.0 of Holmes, which we are now able to offer under a permissive MIT license. Holmes is a library in the spaCy Universe that runs on top of spaCy and enables information extraction and intelligent search, currently for English and German.

Compact word vectors with Bloom embeddings

An introduction to the compact word vectors with Bloom embeddings used in Thinc, spaCy and floret.

Neural edit-tree lemmatization for spaCy

We are happy to introduce a new, experimental, machine learning-based lemmatizer that posts accuracies above 95% for many languages. This lemmatizer learns to predict lemmatization rules from a corpus of examples and removes the need to write an exhaustive set of per-language lemmatization rules.

Applied NLP Thinking: How to Translate Problems into Solutions

We’ve been running Explosion for about five years now, which has given us a lot of insights into what Natural Language Processing looks like in industry contexts. In this blog post, I’m going to discuss some of the biggest challenges for applied NLP and translating business problems into machine learning solutions.

Explosion in 2019: Our Year in Review

As 2019 draws to a close and we step into the 2020s, we thought we’d take a look back at the year and all we’ve accomplished. And we realized we had so much that we could give you a month-by-month rundown of everything that happened.

Training spaCy NER Models with Prodigy

This handy flowchart contains our most common tips, tricks, and best practices for training and updating spaCy named entity recognition models with Prodigy.

How the Guardian approaches quote extraction with NLP

A case study of the Guardian's spaCy-Prodigy workflow to modularize quote extraction for content creation. This study includes iterative annotation guidelines and custom interface functionality.

Introducing spaCy v3.4

spaCy v3.4 brings typing and speed improvements along with new vectors for English CNN pipelines and new trained pipelines for Croatian.

Implementing a custom trainable component for relation extraction

Relation extraction refers to the process of predicting and labeling semantic relationships between named entities. In this blog post, we'll go over the process of building a custom relation extraction component using spaCy and Thinc.

Introducing spaCy v3.5

spaCy v3.5 introduces new CLI commands, fuzzy matching, improvements for entity linking and more.