Explosion builds **developer tools** for AI, Machine Learning and Natural Language Processing. →

[Consulting](/content/tailored-solutions/index.html)

### Project

- [spaCy](/content/site-root.html)
- [Prodigy](/content/_/project/prodigy/index.html)
- [Ellf](/content/_/project/ellf/index.html)
- [Thinc](/content/_/project/thinc/index.html)
- [Consulting](/content/_/project/consulting/index.html)
- [Case Study](/content/_/project/case_study/index.html)

### Topics

- [LLMs](/content/_/topic/llms/index.html)
- [NLP Strategy](/content/_/topic/strategy/index.html)
- [Annotation](/content/_/topic/annotation/index.html)
- [Biomedical](/content/_/topic/biomedical/index.html)
- [Finance](/content/_/topic/finance/index.html)
- [Media](/content/_/topic/media/index.html)
- [Legal](/content/_/topic/legal/index.html)
- [Humanities](/content/_/topic/humanities/index.html)
- [Computer Vision](/content/_/topic/computer-vision/index.html)

### Category

- [Blog](/content/_/category/blog/index.html)
- [Release](/content/_/category/release/index.html)
- [Universe](/content/_/category/universe/index.html)
- [Talk](/content/_/category/talk/index.html)
- [Interview](/content/_/category/interview/index.html)
- [Video](/content/_/category/video/index.html)
- [Paper](/content/_/category/paper/index.html)
- [Book](/content/_/category/book/index.html)

### Tasks

Select...Code GenerationCoreference ResolutionDependency ParsingDistillationEmbeddings & VectorsEntity LinkingEvaluationImage ClassificationImage SegmentationLayout AnalysisLemmatizationNamed Entity RecognitionObject DetectionOptical Character Recognition (OCR)Part-of-Speech TaggingPII AnonymizationQuestion AnsweringRelation ExtractionRetrieval-Augmented Generation (RAG)Rule-Based MatchingSpan CategorizationText ClassificationText GenerationTokenizationWeak Supervision

### Authors

Select...Adriane BoydÁkos KádárBasile DuraChung-Fan TsaiDamian RomeroDaniël de KokDuygu AltinokEdward SchmuhlHelena SteckmeisterIndia KerleInes MontaniKabir KhanLj MirandaMadeesh KannanMagdalena AniołMatthew HonnibalPaul O’Leary McCannPeter BaumgartnerPhilip VolletRaphael MitschRehan AhmedRichard HudsonRyan WesslenSofie Van LandeghemVictoria SlocumVincent D. WarmerdamVinit RavishankarWalter Henry

### Newsletter

Join our mailing list to receive updates about new blog posts and projects!

💌

[View past issues](https://us12.campaign-archive.com/home/?u=83b0498b1e7fa3c91ce68c3f1&id=ecc82e0493)

## [Conquering PDFs: document understanding beyond plain text](https://speakerdeck.com/inesmontani/conquering-pdfs-document-understanding-beyond-plain-text) [PyData London](https://speakerdeck.com/inesmontani/conquering-pdfs-document-understanding-beyond-plain-text)

[In this talk, Ines presents a new and modular approach for building robust document understanding systems, using state-of-the-art models and the awesome Python ecosystem.](https://speakerdeck.com/inesmontani/conquering-pdfs-document-understanding-beyond-plain-text)

[**📚 spacy-layout v0.0.12Mar 8, 2025** \\
Support processing PDFs with context, add document index tables and more docs](https://github.com/explosion/spacy-layout)

## [Mastering spaCy](https://www.packtpub.com/en-us/product/mastering-spacy-9781835880463) [Déborah Mesquita, Duygu Altinok (Packt Publishing, 2025)](https://www.packtpub.com/en-us/product/mastering-spacy-9781835880463)

[Build structured NLP solutions with custom components and models powered by LLMs. By end of the book you will be empowered to build robust NLP pipelines and integrate them with web applications to build end-to-end solutions.](https://www.packtpub.com/en-us/product/mastering-spacy-9781835880463)

## [spaCy Natural Language Processing: From Beginner to Advanced](https://www.linkedin.com/feed/update/urn:li:activity:7274542396934119425/) [Guan Wang, Xiaoquan Kong (2024)](https://www.linkedin.com/feed/update/urn:li:activity:7274542396934119425/)

[The first Chinese-language book on spaCy for beginners and experienced practitioners, covering traditional NLP techniques and how to leverage LLMs for various NLP tasks.](https://www.linkedin.com/feed/update/urn:li:activity:7274542396934119425/)

## [Taking LLMs out of the black box: A practical guide to human-in-the-loop distillation](https://speakerdeck.com/inesmontani/taking-llms-out-of-the-black-box-a-practical-guide-to-human-in-the-loop-distillation) [InfoQ Dev Summit](https://speakerdeck.com/inesmontani/taking-llms-out-of-the-black-box-a-practical-guide-to-human-in-the-loop-distillation)

[LLMs have enormous potential, but also challenge existing workflows in industry that require modularity, transparency and data privacy. In this talk, Ines shows some practical solutions for using the latest models in real-world applications and distilling their knowledge into smaller and faster components that you can run and maintain in-house.](https://speakerdeck.com/inesmontani/taking-llms-out-of-the-black-box-a-practical-guide-to-human-in-the-loop-distillation)

## [Assessing Fine-Tuned NER Models with Limited Data in French: Automating Detection of New Technologies, Technological Domains, and Startup Names in Renewable Energy](https://www.mdpi.com/2504-4990/6/3/96) [MacLean, Cavallucci (2024)](https://www.mdpi.com/2504-4990/6/3/96)

[In order to assure the uniformity of the process of fine-tuning each model, we decided to use the spaCy library. This library, one of the most widely used for NLP tasks, allows us to directly modify a simple configuration file in order to define the model.](https://www.mdpi.com/2504-4990/6/3/96)

## [Back to our roots: Company update and future plans](/content/blog/back-to-our-roots-company-update/index.html)

[We’re back to running Explosion as a smaller, independent-minded and self-sufficient company. spaCy and Prodigy will stay stable and sustainable, maintained by their original authors. We’ll keep updating our stack wth the latest technologies, without changing its core identity or purpose.](/content/blog/back-to-our-roots-company-update/index.html)

## [How S&P Global is making markets more transparent with NLP, spaCy and Prodigy](/content/blog/sp-global-commodities/index.html)

[A case study on S&P Global’s efficient information extraction pipelines for real-time commodities trading insights in a high-security environment.](/content/blog/sp-global-commodities/index.html)

## [spaCy meets LLMs: Using Generative AI for Structured Data](https://speakerdeck.com/inesmontani/spacy-meets-llms-using-generative-ai-for-structured-data) [Data+ML Community Meetup](https://speakerdeck.com/inesmontani/spacy-meets-llms-using-generative-ai-for-structured-data)

[This talk dives deeper into spaCy’s LLM integration, which provides a robust framework for extracting structured information from text, distilling large models into smaller components, and closing the gap between prototype and production.](https://speakerdeck.com/inesmontani/spacy-meets-llms-using-generative-ai-for-structured-data)

## [Constructing a knowledge base with spaCy and spacy-llm](https://medium.com/mantisnlp/constructing-a-knowledge-base-with-spacy-and-spacy-llm-f65b50ea534d) [MantisNLP Blog](https://medium.com/mantisnlp/constructing-a-knowledge-base-with-spacy-and-spacy-llm-f65b50ea534d)

[This blog post shows how to use spaCy and LLMs to extract entities and relationships from text and quickly tackle the complex problem of constructing a knowledge base graph from a corpus.](https://medium.com/mantisnlp/constructing-a-knowledge-base-with-spacy-and-spacy-llm-f65b50ea534d)

## [spacy-llm: From quick prototyping with LLMs to more reliable and efficient NLP solutions](https://speakerdeck.com/sofievl/2024-01-23-az) [AstraZeneca NLP Community of Practice](https://speakerdeck.com/sofievl/2024-01-23-az)

[LLMs are paving the way for fast prototyping of NLP applications. Here, Sofie showcases how to build a structured NLP pipeline to mine clinical trials, using spaCy and spacy-llm. Moving beyond a fast prototype, she offers pragmatic solutions to make the pipeline more reliable and cost efficient.](https://speakerdeck.com/sofievl/2024-01-23-az)

[**🦙 spacy-llm v0.7.0Jan 19, 2024** \\
Supporting arbitrarily long docs and various new tasks](https://github.com/explosion/spacy-llm/releases/tag/v0.7.0)

## [On the Creation of Classifiers to Support Assessment of E-Portfolios](https://aisop.de/publikationen/) [Gantikow, Isking, Libbrecht, Müller, Rebholz (2023)](https://aisop.de/publikationen/)

[In this workflow, Prodigy selects and presents text examples that were classified with a very low degree of certainty. The annotator reviews the proposed classifications and corrects them, if necessary.](https://aisop.de/publikationen/)

## [Developing a Named Entity Recognition Dataset for Tagalog](https://arxiv.org/abs/2311.07161) [Miranda (2023), IJCNLP-AACL 2023](https://arxiv.org/abs/2311.07161)

[We used Prodigy as our annotation tool. We set up a web server on the Google Cloud Platform and routed the examples through Prodigy’s built-in task router.](https://arxiv.org/abs/2311.07161)

## [MP Interests Tracker: Utilising GenAI to uncover insights in the UK Register of Financial Interest](https://www.journalismai.info/blog/mp-interests-tracker-utilising-genai-for-uncovering-insights-in-the-uk-register-of-financial-interest) [JournalismAI Blog](https://www.journalismai.info/blog/mp-interests-tracker-utilising-genai-for-uncovering-insights-in-the-uk-register-of-financial-interest)

[Project from teams at The Times and BBC using spacy-llm to make complex financial interests data more accessible.](https://www.journalismai.info/blog/mp-interests-tracker-utilising-genai-for-uncovering-insights-in-the-uk-register-of-financial-interest)

[**🦙 spacy-llm v0.4.0Jul 6, 2023** \\
Falcon, sentiment analysis, summarization, backend refactoring](https://github.com/explosion/spacy-llm/releases/tag/v0.4.0)

## [Large Language Models: From Prototype to Production](https://speakerdeck.com/inesmontani/large-language-models-from-prototype-to-production) [PyData London Keynote](https://speakerdeck.com/inesmontani/large-language-models-from-prototype-to-production)

## [Implementing a custom trainable component for relation extraction](/content/blog/relation-extraction/index.html)

[Relation extraction refers to the process of predicting and labeling semantic relationships between named entities. In this blog post, we'll go over the process of building a custom relation extraction component using spaCy and Thinc. We'll also add a Hugging Face transformer to improve performance at the end of the post. You'll see how you can utilize Thinc's flexible and customizable system to build an NLP pipeline for biomedical relation extraction.](/content/blog/relation-extraction/index.html)

## [The Tale of Bloom Embeddings and Unseen Entities](/content/blog/technical-report/index.html)

[The default Bloom embedding layer in spaCy is unconventional, but very powerful and efficient. We wrote about it before and showed the advantages it provides in terms of memory efficiency for our floret embeddings. Now we have released the first technical report by Explosion, where we explain Bloom embeddings in more detail and rigorously compare them to traditional embeddings. In this post we'll highlight some of our results with a special focus on unseen entities.](/content/blog/technical-report/index.html)

## [Towards a Tagalog NLP pipeline](https://ljvmiranda921.github.io/notebook/2023/02/04/tagalog-pipeline/)

[In this blog post, Lj talks about how he built an NER pipeline for Tagalog, the gold-standard dataset, benchmarking results, and his hopes for the future of Tagalog NLP.](https://ljvmiranda921.github.io/notebook/2023/02/04/tagalog-pipeline/)

## [Reflections on a year of spaCy consulting at Explosion](https://www.linkedin.com/pulse/reflections-year-spacy-consulting-explosion-peter-baumgartner/)

[In this post, Peter shares some lessons learned from chatting with practitioners about their NLP challenges, developing production-ready NLP pipelines for clients, and working with an open-source development team.](https://www.linkedin.com/pulse/reflections-year-spacy-consulting-explosion-peter-baumgartner/)

## [Is it possible to have entities within entities within entities?](https://speakerdeck.com/victorialslocum/pydata-global-2022-spancat) [PyData Global 2022](https://speakerdeck.com/victorialslocum/pydata-global-2022-spancat)

[Named entity recognition models might not be able to handle a wide variety of spans, but Spancat certainly can! Dive into named entity recognition, its limitations, and how we’ve solved them with a solution-focused talk and practical applications.](https://speakerdeck.com/victorialslocum/pydata-global-2022-spancat)

## [End-to-end Neural Coreference Resolution in spaCy](/content/blog/coref/index.html)

[Coreference resolution is the problem of resolving entities in texts to references such as pronouns. Even if you've never heard of it, it's something we all do constantly every day, and is a key to understanding natural language. We recently added an experimental implementation of an end-to-end neural coreference component to spaCy. This post explains the architecture of our model in detail.](/content/blog/coref/index.html)

[**🧪 spacy-experimental v0.6.0Sep 28, 2022** \\
Added Coref components and models](https://github.com/explosion/spacy-experimental/releases/tag/v0.6.0)

## [Introducing Span Categorization in Prodigy and spaCy](https://www.youtube.com/watch?v=xgV3Rlj49lQ)

[In this video, we’ll show you how to use Prodigy for spaCy’s Span Categorizer. We’ll be annotating food recipes and looking into ways to help with consistent annotations and speed up the process with patterns and temporary models.](https://www.youtube.com/watch?v=xgV3Rlj49lQ)

## [Compact word vectors with Bloom embeddings](/content/blog/bloom-embeddings/index.html)

[An introduction to the compact word vectors with Bloom embeddings used in Thinc, spaCy and floret.](/content/blog/bloom-embeddings/index.html)

## [Explosion in 2021: Our Year in Review](/content/blog/year-in-review-2021/index.html)

[The year 2021 is coming to an end, and like the previous year, it was shaped by unique challenges that impacted our work together. For Explosion, it was a very productive year. We found an investor that fits our strategy, the work on Prodigy Teams is in full swing, and the team has grown a lot. So here's our look back at our highlights of the year 2021.](/content/blog/year-in-review-2021/index.html)

## [spaCy v3's project and config systems are pretty great](/content/blog/spacy-v3-project-config-systems/index.html)

[The road to production has become increasingly harder. Machine Learning\\
Engineers who turn prototypes into production-ready software face\\
difficulties with the lack of tooling and best-practices. spaCy v3, with its\\
configuration and project system, introduced a way to solve this problem. Here's\\
my take on how it works, and how it can ramp-up your team!](/content/blog/spacy-v3-project-config-systems/index.html)

## [Mastering spaCy](https://www.amazon.de/-/en/Mastering-spaCy-end-end-implementing/dp/1800563353) [Duygu Altinok (Packt Publishing, 2021)](https://www.amazon.de/-/en/Mastering-spaCy-end-end-implementing/dp/1800563353)

[An end-to-end practical guide to implementing NLP applications using the Python ecosystem. By the end of this book, you'll be able to confidently use spaCy, including its linguistic features, word vectors, and classifiers, to create your own NLP apps.](https://www.amazon.de/-/en/Mastering-spaCy-end-end-implementing/dp/1800563353)

[**🤗 spacy-huggingface-hub v0.0.1Jul 6, 2021** \\
Upload spaCy pipelines to the Hugging Face Hub](https://github.com/explosion/spacy-huggingface-hub)

## [spaCy v3: Design concepts explained (behind the scenes)](https://www.youtube.com/watch?v=BWhh3r6W-qE)

[In this video, Ines shows you some of the new design concepts and explain what’s going on under the hood, how we’ve implemented them and most importantly, why.](https://www.youtube.com/watch?v=BWhh3r6W-qE)

## [Introducing spaCy v2.3](/content/blog/spacy-v2-3/index.html)

[spaCy now speaks Chinese, Japanese, Danish, Polish and Romanian! Version 2.3 of the spaCy Natural Language Processing library adds models for five new languages. We've also updated all 15 model families with word vectors and improved accuracy, while also decreasing model size and loading times for models with vectors.](/content/blog/spacy-v2-3/index.html)

## [Explosion in 2019: Our Year in Review](/content/blog/year-in-review-2019/index.html)

[As 2019 draws to a close and we step into the 2020s, we thought we’d take a look back at the year and all we’ve accomplished. And we realized we had so much that we could give you a month-by-month rundown of everything that happened.](/content/blog/year-in-review-2019/index.html)

[**🦆 sense2vec v1.0.0Nov 22, 2019** \\
More features, 2019 Reddit vectors model and Prodigy recipes](https://github.com/explosion/sense2vec/releases/tag/v1.0.0)

## [Explosion awarded META Seal of Recognition](http://www.meta-net.eu/meta-seal)

[We’re proud to accept the META Seal of Recognition at META-FORUM in Brussels, along with Mozilla. The META-FORUM is an international conference series backed by the European Union on powerful and innovative Language Technologies for a multilingual information society.](http://www.meta-net.eu/meta-seal)

## [Intro to NLP with spaCy (1): Detecting programming languages](https://www.youtube.com/watch?v=WnGPv6HnBok&list=PLBmcuObd5An559HbDr_alBnwVsGq-7uTF&index=1)

[In this new video series, data science instructor Vincent Warmerdam gets started with spaCy, an open-source library for Natural Language Processing in Python. His mission: building a system to automatically detect programming languages in large volumes of text.](https://www.youtube.com/watch?v=WnGPv6HnBok&list=PLBmcuObd5An559HbDr_alBnwVsGq-7uTF&index=1)

## [McKenzie Marshall: NLP in Asset Management (Barings)](https://www.youtube.com/watch?v=kX14Ycieju8&list=PLBmcuObd5An4UC6jvK_-eSl6jCvP1gwXc&index=10) [spaCy IRL 2019](https://www.youtube.com/watch?v=kX14Ycieju8&list=PLBmcuObd5An4UC6jvK_-eSl6jCvP1gwXc&index=10)

## [spaCy IRL 2019: 2 days of NLP in Berlin](https://irl.spacy.io/2019)

[We were pleased to invite the spaCy community and other folks working on Natural Language Processing to Berlin this summer for a small and intimate event.](https://irl.spacy.io/2019)

## [Practical transfer learning for NLP with spaCy and Prodigy](https://speakerdeck.com/inesmontani/practical-transfer-learning-for-nlp-with-spacy-and-prodigy) [Applied Machine Learning Days](https://speakerdeck.com/inesmontani/practical-transfer-learning-for-nlp-with-spacy-and-prodigy)

## [Explosion in 2017: Our Year in Review](/content/blog/year-in-review-2017/index.html)

[We founded Explosion in October 2016, so this was our first full calendar year in operation. We set ourselves ambitious goals this year, and we're very happy with how we achieved them. Here's what we got done.](/content/blog/year-in-review-2017/index.html)

## [Training an insults classifier with Prodigy in ~1 hour](https://www.youtube.com/watch?v=5di0KlKl0fE)

[In this video, we’ll show you how to use Prodigy to train a classifier to detect disparaging or insulting comments. Prodigy makes text classification particularly powerful, because you can try out new ideas very quickly.](https://www.youtube.com/watch?v=5di0KlKl0fE)

## [spaCy now speaks German](/content/blog/german-model/index.html)

[Many people have asked us to make spaCy available for their language. Being based in Berlin, German was an obvious choice for our first second language. Now spaCy can do all the cool things you use for processing English on German text too. But more importantly, teaching spaCy to speak German required us to drop some comfortable but English-specific assumptions about how language works and made spaCy fit to learn more languages in the future.](/content/blog/german-model/index.html)

## [E^2GraphRAG: Streamlining Graph-based RAG for High Efficiency and Effectiveness](https://arxiv.org/abs/2505.24226) [Zhao, Zhu, Guo, He, Li (2025)](https://arxiv.org/abs/2505.24226)

[Instead of using LLMs for entity extraction, we employ the traditional NLP tool spaCy to extract entities, and use their co-occurrence in a chunk as relations.](https://arxiv.org/abs/2505.24226)

## [Prozessvisualisierung mit generativer KI im Praxistest](https://www.heise.de/ratgeber/Prozessvisualisierung-mit-generativer-KI-im-Praxistest-10266093.html) [iX Magazin / Heise](https://www.heise.de/ratgeber/Prozessvisualisierung-mit-generativer-KI-im-Praxistest-10266093.html)

[German article by Nils Durner on visualizing technical processes with Generative AI, featuring spaCy and Presidio for PII anonymization.](https://www.heise.de/ratgeber/Prozessvisualisierung-mit-generativer-KI-im-Praxistest-10266093.html)

## [From PDFs to AI-ready structured data: a deep dive](/content/blog/pdfs-nlp-structured-data/index.html)

[This blog post presents a new modular workflow for converting PDFs and similar documents to structured data and shows you how to build end-to-end document understanding and information extraction pipelines for industry use cases.](/content/blog/pdfs-nlp-structured-data/index.html)

[**📚 spacy-layout v0.0.6Nov 24, 2024** \\
Add support for tables and convert tabular data to pandas.DataFrame](https://github.com/explosion/spacy-layout)

## [Applied NLP in the Age of Generative AI](https://speakerdeck.com/inesmontani/applied-nlp-in-the-age-of-generative-ai) [PyData Amsterdam Keynote](https://speakerdeck.com/inesmontani/applied-nlp-in-the-age-of-generative-ai)

[In this talk, Ines shares the most important lessons we’ve learned from solving real-world information extraction problems in industry, and shows you a new approach and mindset for designing robust and modular NLP pipelines in the age of Generative AI.](https://speakerdeck.com/inesmontani/applied-nlp-in-the-age-of-generative-ai)

## [The NLP and AI Revolution with the spaCy Creators](https://vanishinggradients.fireside.fm/34) [Vanishing Gradients](https://vanishinggradients.fireside.fm/34)

[In this interview with Hugo Bowne-Anderson, we delve into the forefront of NLP and the future of AI development, covering topics like human-in-the-loop distillation, open-source AI and Explosion’s journey.](https://vanishinggradients.fireside.fm/34)

## [Building the Future of NLP: Insights on spaCy, Prodigy and Generative AI](https://community.analyticsvidhya.com/c/leading-with-data/ines) [Leading With Data Podcast](https://community.analyticsvidhya.com/c/leading-with-data/ines)

## [Exploring the AI nexus with the mind behind spaCy](https://community.analyticsvidhya.com/c/leading-with-data/matthew-honnibal) [Leading With Data Podcast](https://community.analyticsvidhya.com/c/leading-with-data/matthew-honnibal)

[In this episode, Matt takes you on a deep dive into the future of data and the challenges facing current Large Language Models (LLMs).](https://community.analyticsvidhya.com/c/leading-with-data/matthew-honnibal)

## [Getting Started with NLP and spaCy](https://training.talkpython.fm/courses/getting-started-with-spacy) [TalkPython Course](https://training.talkpython.fm/courses/getting-started-with-spacy)

[There is a lot of text data out there and maybe you're interested in getting structured data out of it. There are a lot of options out there and this course will introduce you to the field by focussing on spaCy while also exploring other tools.](https://training.talkpython.fm/courses/getting-started-with-spacy)

## [T-RAG: Lessons from the LLM Trenches](https://arxiv.org/abs/2402.07483) [Fatehkia, Lucas, Chawla (2024)](https://arxiv.org/abs/2402.07483)

[An important application area is question answering over private enterprise documents where the main considerations are data security, which necessitates applications that can be deployed on-prem, \[and\] limited computational resources. \[...\] In addition to retrieving contextual documents, we use the spaCy library with custom rules to detect named entities from the organization.](https://arxiv.org/abs/2402.07483)

## [Microsoft Presidio v2.2.352](https://github.com/microsoft/presidio)

[Context aware, pluggable and customizable PII de-identification and anonymization service for text and images, featuring a spaCy back-end.](https://github.com/microsoft/presidio)

## [Neuradicon: operational representation learning of neuroimaging reports](https://arxiv.org/abs/2107.10021) [Watkins, Gray, Julius, Mah, Pinaya, Wright, Jha, Engleitner, Cardoso, Ourselin, Rees, Jaeger, Nachev (2023)](https://arxiv.org/abs/2107.10021)

[Labelled data for each task was produced using the Prodigy labelling tool. Each report was labelled in a paired-annotation manner. \[...\] We used the grammatical dependency parse produced by the spaCy parser as input and implemented the patterns using the spaCy dependency matcher.](https://arxiv.org/abs/2107.10021)

## [State-of-the-Art Transformer Pipelines in spaCy](https://www.youtube.com/watch?v=clnhMaTq1ZA) [aiGrunn](https://www.youtube.com/watch?v=clnhMaTq1ZA)

[In this talk, we will show you how you can use transformer models (from pretrained models such as XLM-RoBERTa to large language models like Llama2) to create state-of-the-art annotation pipelines for text annotation tasks such as named entity recognition.](https://www.youtube.com/watch?v=clnhMaTq1ZA)

[**🦙 spacy-llm v0.6.0Oct 5, 2023** \\
PaLM, Azure OpenAI, Mistral & fixed OS model responses](https://github.com/explosion/spacy-llm/releases/tag/v0.6.0)

## [Large Language Models: From Prototype to Production](https://speakerdeck.com/inesmontani/large-language-models-from-prototype-to-production-europython-keynote) [EuroPython Keynote](https://speakerdeck.com/inesmontani/large-language-models-from-prototype-to-production-europython-keynote)

[Large Language Models (LLMs) have shown some impressive capabilities and their impact is the topic of the moment. In this talk, Ines presents visions for NLP in the age of LLMs and a pragmatic, practical approach for how to use Large Language Models to ship more successful NLP projects from prototype to production today.](https://speakerdeck.com/inesmontani/large-language-models-from-prototype-to-production-europython-keynote)

[**🦙 spacy-llm v0.3.0Jun 14, 2023** \\
Cohere, Anthropic, OpenLLaMa, StableLM, logging, streamlit demo, lemmatization task](https://github.com/explosion/spacy-llm/releases/tag/v0.3.0)

## [SpanCat with spaCy and Prodigy on real data](https://www.youtube.com/watch?v=6S52SUBFZxc&list=PL2VXyKi-KpYtKSdydjcsI3L8dUj4Ck3iP)

[YouTube series by WJB Mattingly showing an end-to-end project, from cultivating and annotating data to training, testing and visualizing a model.](https://www.youtube.com/watch?v=6S52SUBFZxc&list=PL2VXyKi-KpYtKSdydjcsI3L8dUj4Ck3iP)

## [You are what you read: Building a personal internet front-page with spaCy and Prodigy](https://speakerdeck.com/victorialslocum/pydata-berlin-2023) [PyCon DE & PyData Berlin](https://speakerdeck.com/victorialslocum/pydata-berlin-2023)

## [Creating Custom Event Data Without Dictionaries: A Bag-of-Tricks](https://arxiv.org/abs/2304.01331) [Halterman, Schrodt, Beger, Bagozzi, Scarborough (2023)](https://arxiv.org/abs/2304.01331)

[While in the past the process of generating training case has been quite time consuming and tedious, newer approaches such as those incorporated into the web-based Prodigy annotation system allow this to be done much more quickly.](https://arxiv.org/abs/2304.01331)

## [Introducing spaCy v3.5](/content/blog/spacy-v3-5/index.html)

[spaCy v3.5 introduces new CLI commands, fuzzy matching, improvements for entity linking and more.](/content/blog/spacy-v3-5/index.html)

## [WW2 spaCy v0.0.9](https://github.com/wjbmattingly/ww2-spacy)

[spaCy pipeline for processing primary and secondary sources for World War 2 texts.](https://github.com/wjbmattingly/ww2-spacy)

## [Coreference Resolution in spaCy](https://www.youtube.com/watch?v=fio3BejnRsM)

[In everyday conversation, we use pronouns or other expressions to refer to entities in many different ways, but we effortlessly understand these references. In NLP this is a challenging problem known as Coreference Resolution. In this video, we’ll show how to train spaCy’s new component for Coreference Resolution and how to apply the pipeline to resolve references in a text.](https://www.youtube.com/watch?v=fio3BejnRsM)

## [spaCy behind the scenes: library patterns & design concepts explained](/content/blog/spacy-design-concepts/index.html)

[Developer productivity has been central to our design of spaCy, both in smaller decisions and some of the bigger architectural questions. We believe in embracing the complexities of machine learning, not hiding it away under leaky abstractions, while also maintaining the developer experience. Read on to learn some of the design patterns within the library, how we've implemented them, and most importantly, why.](/content/blog/spacy-design-concepts/index.html)

## [Spancat: a new approach for span labeling](/content/blog/spancat/index.html)

[The SpanCategorizer is a spaCy component that answers the NLP community's need to\\
have structured annotation for a wide variety of labeled spans, including long\\
phrases, non-named entities, or overlapping annotations. In this blog post, we're\\
excited to talk more about spancat and showcase new features to help with your\\
span labeling needs!](/content/blog/spancat/index.html)

[**🧪 spacy-experimental v0.5.0Jun 11, 2022** \\
Added SpanFinder, Span suggesters and bugfixes](https://github.com/explosion/spacy-experimental/releases/tag/v0.5.0)

## [skweak v0.3.1](https://github.com/NorskRegnesentral/skweak)

[Weak supervision and flexible label functions and agrregation, integrated with spaCy.](https://github.com/NorskRegnesentral/skweak)

## [Healthsea: an end-to-end spaCy pipeline for exploring health supplement effects](/content/blog/healthsea/index.html)

[Create better access to health with machine learning and natural language processing. Read about our journey of developing Healthsea, an end-to-end spaCy pipeline for analyzing user reviews to supplement products and extracting potential effects on health.](/content/blog/healthsea/index.html)

## [Introducing spaCy v3.2](/content/blog/spacy-v3-2/index.html)

[spaCy v3.2 features usability improvements for custom training and scoring, improved performance and support for floret, our new fastText word vectors algorithm.](/content/blog/spacy-v3-2/index.html)

## [Introducing spaCy v3.1](/content/blog/spacy-v3-1/index.html)

[It’s been great to see the adoption of spaCy v3, which introduced transformer-based pipelines, a new training system and more. Version 3.1 adds more on top of it, including the ability to use predicted annotations during training, a component for predicting arbitrary and overlapping spans and new pipelines for Catalan and Danish.](/content/blog/spacy-v3-1/index.html)

[**🦆 sense2vec v2.0.0Feb 7, 2021** \\
Update component for spaCy v3](https://github.com/explosion/sense2vec/releases/tag/v2.0.0)

## [spaCy v3: State-of-the-art NLP from Prototype to Production](https://www.youtube.com/watch?v=9k_EfV7Cns0)

## [Building customizable NLP pipelines with spaCy](https://speakerdeck.com/sofievl/2020-02-19-spacy-pipelines) [Turku.AI Meetup](https://speakerdeck.com/sofievl/2020-02-19-spacy-pipelines)

## [Intro to NLP with spaCy (3): Detecting programming languages](https://www.youtube.com/watch?v=4V0JDdohxAk&list=PLBmcuObd5An559HbDr_alBnwVsGq-7uTF&index=3)

## [Introducing spaCy v2.2](/content/blog/spacy-v2-2/index.html)

[Version 2.2 of the spaCy Natural Language Processing library is leaner, cleaner and even more user-friendly. In addition to new model packages and features for training, evaluation and serialization, we've made lots of bug fixes, improved debugging and error handling, and greatly reduced the size of the library on disk.](/content/blog/spacy-v2-2/index.html)

## [Blackstone v0.1.15](https://github.com/ICLRandD/Blackstone)

[A spaCy pipeline and model for NLP on unstructured legal text](https://github.com/ICLRandD/Blackstone)

## [Patrick Harrison: Financial NLP at S&P Global](https://www.youtube.com/watch?v=rdmaR4WRYEM&list=PLBmcuObd5An4UC6jvK_-eSl6jCvP1gwXc&index=9) [spaCy IRL 2019](https://www.youtube.com/watch?v=rdmaR4WRYEM&list=PLBmcuObd5An4UC6jvK_-eSl6jCvP1gwXc&index=9)

## [Practical transfer learning for NLP with spaCy and Prodigy](https://www.youtube.com/watch?v=dkJnI70mTk4&ab_channel=Infoshare) [Infoshare](https://www.youtube.com/watch?v=dkJnI70mTk4&ab_channel=Infoshare)

## [The process: Transforming spaCy’s docs](https://increment.com/documentation/transforming-spacys-docs/) [Increment Magazine](https://increment.com/documentation/transforming-spacys-docs/)

[Making your documentation work for users with vastly different needs is a challenge. Here’s how spaCy, an open-source library for natural language processing, did it.](https://increment.com/documentation/transforming-spacys-docs/)

## [Training a new entity type with Prodigy – annotation powered by active learning](https://www.youtube.com/watch?v=l4scwf8KeIA)

[In this video, we’ll show you how to use Prodigy to train a phrase recognition system for a new concept. Specifically, we’ll train a model to detect references to drugs, using text from Reddit.](https://www.youtube.com/watch?v=l4scwf8KeIA)

## [Pseudo-rehearsal: A simple solution to catastrophic forgetting for NLP](/content/blog/pseudo-rehearsal-catastrophic-forgetting/index.html)

[Sometimes you want to fine-tune a pre-trained model to add a new label or correct some specific errors. This can introduce the "catastrophic forgetting" problem. Pseudo-rehearsal is a good solution: use the original model to label examples, and mix them through your fine-tuning updates.](/content/blog/pseudo-rehearsal-catastrophic-forgetting/index.html)

## [Sense2vec with spaCy and Gensim](/content/blog/sense2vec-with-spacy/index.html)

[If you were doing text analytics in 2015, you were probably using word2vec. Sense2vec (Trask et. al, 2015) is a new twist on word2vec that lets you learn more interesting, detailed and context-sensitive word vectors. This post motivates the idea, explains our implementation, and comes with an interactive demo that we've found surprisingly addictive.](/content/blog/sense2vec-with-spacy/index.html)

## [Conquering PDFs: document understanding beyond plain text](https://speakerdeck.com/inesmontani/conquering-pdfs-document-understanding-beyond-plain-text) [PyCon DE & PyData](https://speakerdeck.com/inesmontani/conquering-pdfs-document-understanding-beyond-plain-text)

## [Using natural language processing to identify emergency department patients with incidental lung nodules requiring follow-up](https://onlinelibrary.wiley.com/doi/abs/10.1111/acem.15080) [Moore, Socrates, Hesami, Denkewicz, Cavallo, Venkatesh, Taylor (2025)](https://onlinelibrary.wiley.com/doi/abs/10.1111/acem.15080)

[CT reports were annotated by MD raters using Prodigy software to develop a stepwise NLP “pipeline” that first excluded prior or known malignancy, determined the presence of a lung nodule, and then categorized any recommended follow-up. NLP was developed using a RoBERTa large language model on the spaCy platform.](https://onlinelibrary.wiley.com/doi/abs/10.1111/acem.15080)

[**📚 spacy-layout v0.0.1Nov 18, 2024** \\
Process PDFs, Word documents and more with spaCy](https://github.com/explosion/spacy-layout)

## [uOttawa at LegalLens-2024: Transformer-based Classification Experiments](https://arxiv.org/abs/2410.21139) [Meghdadi, Inkpen (2024)](https://arxiv.org/abs/2410.21139)

[Our training utilizes the spaCy pipeline configured with a transformer model and a transition-based parser for NER tasks. The deberta-v3-base model has been selected for the main transformer architecture.](https://arxiv.org/abs/2410.21139)

## [Combining the Best of Two Worlds: From TF-IDF to Llama LLM](https://osseu2024.sched.com/event/b3c1139bb641e25f16be9451b5123365) [Open Source Summit Europe](https://osseu2024.sched.com/event/b3c1139bb641e25f16be9451b5123365)

[Talk by William Arias, Staff Developer Advocate at GitLab, on combining traditional NLP techniques and LLMs to solve hallucination issues and create robust spaCy applications.](https://osseu2024.sched.com/event/b3c1139bb641e25f16be9451b5123365)

## [spaCy Chunks v0.0.2](https://github.com/wjbmattingly/spacy-chunks)

[spaCy extension and pipeline component for generating overlapping chunks of sentences or tokens from a document.](https://github.com/wjbmattingly/spacy-chunks)

## [Happy 10th Birthday, spaCy!](https://www.linkedin.com/feed/update/urn:li:activity:7214245844407988225/)

[10 years ago today Matt pushed the first commit to spaCy. Since then, the library has evolved as the field moved forward, but also stayed true to its core mission: industrial-strength NLP.](https://www.linkedin.com/feed/update/urn:li:activity:7214245844407988225/)

## [Taking LLMs out of the black box: A practical guide to human-in-the-loop distillation](https://speakerdeck.com/inesmontani/taking-llms-out-of-the-black-box-a-practical-guide-to-human-in-the-loop-distillation) [PyData London](https://speakerdeck.com/inesmontani/taking-llms-out-of-the-black-box-a-practical-guide-to-human-in-the-loop-distillation)

## [The application of natural language processing for the extraction of mechanistic information in toxicology](https://www.frontiersin.org/journals/toxicology/articles/10.3389/ftox.2024.1393662/full) [Conradi, Luechtefeld, de Haan, Pieters, Freedman, Vanhaecke, Vinken, Teunis (2024)](https://www.frontiersin.org/journals/toxicology/articles/10.3389/ftox.2024.1393662/full)

[All steps were conducted using the open-source Python package spaCy. Specifically, the NER model was trained using scispaCy en-core-sci-lg (Neumann et al., 2019) as a starting point, which allowed for a vocabulary (word vectors) and grammar trained on scientific literature.](https://www.frontiersin.org/journals/toxicology/articles/10.3389/ftox.2024.1393662/full)

## [How Nesta uses NLP to process 7m job ads and shed light on the UK’s labor market](/content/blog/nesta-skills/index.html)

[A case study on Nesta’s workflow for extracting 7 million job ads to better understand UK skill demand, using a custom mapping step to match skills to any government taxonomy.](/content/blog/nesta-skills/index.html)

## [Muted: Multilingual Targeted Offensive Speech Identification and Visualization](https://arxiv.org/abs/2312.11344) [Tillmann, Trivedi, Rosenthal, Borse, Zhang, Sil, Bhattacharjee (2023)](https://arxiv.org/abs/2312.11344)

[Muted can leverage any transformer-based HAP-classification model \[...\] to identify toxic spans, without further fine-tuning. In addition, we use the spaCy library to identify the specific targets and arguments for the words predicted by the attention heatmaps.](https://arxiv.org/abs/2312.11344)

## [Who said what: using machine learning to correctly attribute quotes](https://www.theguardian.com/info/2023/nov/21/who-said-what-using-machine-learning-to-correctly-attribute-quotes) [The Guardian Engineering Blog](https://www.theguardian.com/info/2023/nov/21/who-said-what-using-machine-learning-to-correctly-attribute-quotes)

[How the Guardian uses spaCy and Prodigy to train a custom coreference resolution model.](https://www.theguardian.com/info/2023/nov/21/who-said-what-using-machine-learning-to-correctly-attribute-quotes)

## [GERNERMED++: Semantic annotation in German medical NLP through transfer-learning, translation and word alignment](https://www.sciencedirect.com/science/article/pii/S1532046423002344) [Frei, Frei-Stuber, Kramer (2023), Journal of Biomedical Informatics](https://www.sciencedirect.com/science/article/pii/S1532046423002344)

[The training of our entity recognition model employs the entity recognition parser from the spaCy library which follows a transducer-based parsing approach with a BILOU scheme instead of a state-agnostic token tagging approach.](https://www.sciencedirect.com/science/article/pii/S1532046423002344)

[**💫 spacy v3.7.0Oct 2, 2023** \\
Trained pipelines using Curated Transformers and support for Python 3.12](https://github.com/explosion/spaCy/releases/tag/v3.7.0)

## [Introducing spaCy v3.6](/content/blog/spacy-v3-6/index.html)

[spaCy v3.6 introduces the span finder component and trained pipelines for Slovenian.](/content/blog/spacy-v3-6/index.html)

[**🦙 spacy-llm v0.2.0May 30, 2023** \\
REL and spancat tasks, reading prompt templates from file](https://github.com/explosion/spacy-llm/releases/tag/v0.2.0)

## [Large Disagreement Modelling](https://koaning.io/posts/large-disagreement-models/)

[“In this blogpost I’d like to talk about large language models. There’s a bunch of hype, sure, but there’s also an opportunity to revisit one of my favourite machine learning techniques: disagreement.”](https://koaning.io/posts/large-disagreement-models/)

## [Incorporating LLMs into practical NLP workflows](https://speakerdeck.com/inesmontani/incorporating-llms-into-practical-nlp-workflows) [PyCon DE & PyData Berlin](https://speakerdeck.com/inesmontani/incorporating-llms-into-practical-nlp-workflows)

## [textaCy v0.13.0](https://textacy.readthedocs.io/en/latest/)

[Utility library for NLP tasks before and after spaCy, including preprocessing, normalization and additional information extraction features.](https://textacy.readthedocs.io/en/latest/)

## [Explosion in 2022: Our Year in Review](/content/blog/year-in-review-2022/index.html)

[It's been another exciting year at Explosion! We've developed a new end-to-end neural coref component for spaCy, improved the speed of our CNN pipelines up to 60%, and published new pre-trained pipelines for Finnish, Korean, Swedish and Croatian. We've also released several updates to Prodigy and introduced new recipes to kickstart annotation with zero- or few-shot learning.](/content/blog/year-in-review-2022/index.html)

## [Setting your ML project up for success](https://www.linkedin.com/pulse/setting-your-ml-project-up-success-sofie-van-landeghem/)

[“What can you do to maximize probability of success for your Machine Learning solution? Throughout my 15 years as data scientist in academia, big pharma and through consulting, one common theme has emerged: the most reliable predictor of success for any NLP or ML-based solution is whether or not you involve the data science team early on.”](https://www.linkedin.com/pulse/setting-your-ml-project-up-success-sofie-van-landeghem/)

## [medspacy v1.0](https://github.com/medspacy/medspacy)

[A library of tools for performing clinical NLP and text processing tasks with spaCy.](https://github.com/medspacy/medspacy)

## [floret: lightweight, robust word vectors](/content/blog/floret-vectors/index.html)

[An exploration of floret vectors: lightweight vectors for noisy data, novel words, rich morphology and more.](/content/blog/floret-vectors/index.html)

## [Diary of a spaCy project: Predicting GitHub Tags](/content/blog/diary-of-github-spacy-project/index.html)

[Many people assume that working on an NLP project involves a lot of machine learning. Our experience is that it's much less about flowing tensors, and more about making a tailored solution. This blogposts demonstrates how a typical spaCy project could be initiated, implemented and executed towards a custom solution.](/content/blog/diary-of-github-spacy-project/index.html)

## [Applied Language Technology](https://applied-language-technology.mooc.fi/html/index.html)

[Extensive online course on applied language technology with spaCy by Tuomo Hiippala, designed for students new to NLP and programming.](https://applied-language-technology.mooc.fi/html/index.html)

[**🧪 spacy-experimental v0.4.0Mar 22, 2022** \\
Added biaffine parser and other fixes for experimental tools](https://github.com/explosion/spacy-experimental/releases/tag/v0.4.0)

## [Universal Dependencies v2.5 Benchmarks for spaCy](/content/blog/ud-benchmarks-v3-2/index.html)

[We present Universal Dependencies v2.5 benchmarks for spaCy v3.2 that show\\
the competitive performance of spaCy in a direct comparison with Stanza and\\
Trankit using the end-to-end evaluation from the CoNLL 2018 Shared Task.](/content/blog/ud-benchmarks-v3-2/index.html)

## [Reproducible spaCy NLP Experiments with Weights & Biases](https://wandb.ai/wandb/wandb_spacy_integration/reports/Reproducible-spaCy-NLP-Experiments-with-Weights-Biases--Vmlldzo4NjM2MDk) [Weights & Biases Blog](https://wandb.ai/wandb/wandb_spacy_integration/reports/Reproducible-spaCy-NLP-Experiments-with-Weights-Biases--Vmlldzo4NjM2MDk)

[This tutorial will show how to add Weights & Biases to any spaCy NLP project to track your experiments, save model checkpoints, and version your datasets.](https://wandb.ai/wandb/wandb_spacy_integration/reports/Reproducible-spaCy-NLP-Experiments-with-Weights-Biases--Vmlldzo4NjM2MDk)

## [Intro to NLP with spaCy (6): Detecting programming languages](https://www.youtube.com/watch?v=k77RrmMaKEI&list=PLBmcuObd5An559HbDr_alBnwVsGq-7uTF&index=6)

[**🛸 spacy-transformers v1.0.0Feb 1, 2021** \\
Update components for spaCy v3.0](https://github.com/explosion/spacy-transformers/releases/tag/v1.0.0)

## [Introducing spaCy v3.0](/content/blog/spacy-v3/index.html)

[spaCy v3.0 is a huge release! It features new transformer-based pipelines that get spaCy's accuracy right up to the current state-of-the-art, and a new workflow system to help you take projects from prototype to production. It's much easier to configure and train your pipeline, and there are lots of new and improved integrations with the rest of the NLP ecosystem.](/content/blog/spacy-v3/index.html)

## [Intro to NLP with spaCy (5): Detecting programming languages](https://www.youtube.com/watch?v=f4sqeLRzkPg&list=PLBmcuObd5An559HbDr_alBnwVsGq-7uTF&index=5)

## [sense2vec reloaded: contextually-keyed word vectors](/content/blog/sense2vec-reloaded/index.html)

[In 2016 we trained a sense2vec model on the 2015 portion of the Reddit comments corpus, leading to a useful library and one of our most popular demos. That work is now due for an update. In this post, we present a new version and a demo NER project that we trained to usable accuracy in just a few hours.](/content/blog/sense2vec-reloaded/index.html)

## [Entity linking for spaCy: Grounding textual mentions](https://speakerdeck.com/sofievl/2019-10-01-el-spacy) [Belgium NLP Meetup](https://speakerdeck.com/sofievl/2019-10-01-el-spacy)

## [spaCy meets Transformers: Fine-tune BERT, XLNet and GPT-2](/content/blog/spacy-transformers/index.html)

[Huge transformer models like BERT, GPT-2 and XLNet have set a new standard for accuracy on almost every NLP leaderboard. You can now use these models in spaCy, via a new interface library we've developed that connects spaCy to Hugging Face's awesome implementations.](/content/blog/spacy-transformers/index.html)

## [Mark Neumann: ScispaCy: A spaCy pipeline & models for scientific & biomedical text](https://www.youtube.com/watch?v=2_HSKDALwuw&list=PLBmcuObd5An4UC6jvK_-eSl6jCvP1gwXc&index=8) [spaCy IRL 2019](https://www.youtube.com/watch?v=2_HSKDALwuw&list=PLBmcuObd5An4UC6jvK_-eSl6jCvP1gwXc&index=8)

## [Advanced NLP with spaCy: A free online course](https://course.spacy.io/)

[In this free and interactive online course, you’ll learn how to use spaCy to build advanced natural language understanding systems, using both rule-based and machine learning approaches.](https://course.spacy.io/)

## [What 1.2 million parliamentary speeches can teach us about gender representation](https://pudding.cool/2018/07/women-in-parliament/) [The Pudding](https://pudding.cool/2018/07/women-in-parliament/)

[Analysis of parliamentary speeches using spaCy.](https://pudding.cool/2018/07/women-in-parliament/)

## [More than a Million Pro-Repeal Net Neutrality Comments were Likely Faked](https://hackernoon.com/more-than-a-million-pro-repeal-net-neutrality-comments-were-likely-faked-e9f0e3ed36a6) [Hackernoon](https://hackernoon.com/more-than-a-million-pro-repeal-net-neutrality-comments-were-likely-faked-e9f0e3ed36a6)

[Analysis of net neutrality comments by Jeff Kao using spaCy for word vectors.](https://hackernoon.com/more-than-a-million-pro-repeal-net-neutrality-comments-were-likely-faked-e9f0e3ed36a6)

## [Reflections on running spaCy: commercial open-source NLP](https://ines.io/blog/spacy-commercial-open-source-nlp/) [ines.io](https://ines.io/blog/spacy-commercial-open-source-nlp/)

[As more and more people and companies are getting involved with open-source software, balancing the expectations of an open community and a traditional provider vs. consumer relationship is becoming increasingly difficult. Are maintainers becoming too authoritarian? Are users becoming too demanding? Are large companies selling out open-source?](https://ines.io/blog/spacy-commercial-open-source-nlp/)

## [How spaCy Works](/content/blog/how-spacy-works/index.html)

[This post was pushed out in a hurry, immediately after spaCy was released. It explains some of how spaCy is designed and implemented, and provides some quick notes explaining which algorithms were used. The post pre-dates spaCy's named entity recogniser, but it provides some detail about the tokenisation algorithm, general design, and efficiency concerns.](/content/blog/how-spacy-works/index.html)

## [Keyword Extraction, and Aspect Classification in Sinhala, English, and Code-Mixed Content](https://arxiv.org/abs/2504.10679) [Rizvi, Navojith, Adhikari, Senevirathna, Kasthurirathna, Abeywardhana (2025)](https://arxiv.org/abs/2504.10679)

[Keyword extraction in English is performed with a hybrid approach comprising a fine-tuned spaCy NER model, FinBERT-based KeyBERT embeddings, YAKE, and EmbedRank, which results in a combined accuracy of 91.2%.](https://arxiv.org/abs/2504.10679)

## [Best Way to OCR a PDF in Python](https://www.youtube.com/watch?v=quJtzVxoMtE) [Python Tutorials for Digital Humanities](https://www.youtube.com/watch?v=quJtzVxoMtE)

[Tutorial by WJB Mattingly on how to use the new spaCy Layout package and Docling to convert PDFs to text.](https://www.youtube.com/watch?v=quJtzVxoMtE)

## [Serverless custom NLP with LLMs, Modal and Prodigy](/content/blog/modal-prodigy-serverless-nlp/index.html)

[In this blog post, we’ll show you how you can go from an idea and little data to a fully custom information extraction model using Prodigy and Modal, no infrastructure or GPU setup required.](/content/blog/modal-prodigy-serverless-nlp/index.html)

[**💫 spacy v3.8.0Oct 1, 2024** \\
Memory management for persistent services, numpy 2.0 support](https://github.com/explosion/spaCy/releases/tag/release-v3.8.2)

## [How GitLab uses spaCy to analyze support tickets and empower their community](/content/blog/gitlab-support-insights/index.html)

[A case study on GitLab’s large-scale NLP pipelines for extracting actionable insights from support tickets and usage questions.](/content/blog/gitlab-support-insights/index.html)

## [Practical Tips for Bootstrapping Information Extraction Pipelines](https://speakerdeck.com/honnibal/practical-tips-for-bootstrapping-information-extraction-pipelines) [DataHack Summit](https://speakerdeck.com/honnibal/practical-tips-for-bootstrapping-information-extraction-pipelines)

[This talk presents approaches for bootstrapping NLP pipelines and retrieval via information extraction, including tips for training, modelling and data annotation.](https://speakerdeck.com/honnibal/practical-tips-for-bootstrapping-information-extraction-pipelines)

## [Once a Maintainer: Sofie Van Landeghem](https://onceamaintainer.substack.com/p/once-a-maintainer-sofie-van-landeghem)

[Interview with Sofie about her work as a core maintainer of spaCy, the evolution of NLP, and why dependency management in Python is so terrible.](https://onceamaintainer.substack.com/p/once-a-maintainer-sofie-van-landeghem)

## [Simply Simplify Language](https://github.com/machinelearningZH/simply-simplify-language)

[Interactive app by the Canton of Zurich, Switzerland, using LLMs and spaCy to analyze and simplify institutional communication and make bureaucratic German more inclusive.](https://github.com/machinelearningZH/simply-simplify-language)

## [spaCyEx v0.0.2](https://github.com/wjbmattingly/spacyex)

[Extension for spaCy’s powerful, linguistically-aware pattern matching that introduces a RegEx-like syntax.](https://github.com/wjbmattingly/spacyex)

## [Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes](https://arxiv.org/abs/2402.01352) [Takmaz, Pezzelle, Fernández (2024)](https://arxiv.org/abs/2402.01352)

[We use the spaCy library for tokenization, part-of-speech tagging, and lemmatization of the words in the descriptions.](https://arxiv.org/abs/2402.01352)

## [Herding LLMs Towards Structured NLP](https://speakerdeck.com/rmitsch/herding-llms-towards-structured-nlp) [Global AI Conference](https://speakerdeck.com/rmitsch/herding-llms-towards-structured-nlp)

[This talk shows how we integrate LLMs into spaCy, leveraging its modular and customizable framework. This allows for cheaper, faster and more robust NLP - driven by cutting-edge LLMs, without compromising on having structured, validated data.](https://speakerdeck.com/rmitsch/herding-llms-towards-structured-nlp)

## [Launching the Explosion Merch Store](/content/merch/index.html)

[Spread the love and support us and our open-source work with some of our unique, custom-designed swag. All orders come with free shipping and stickers!](/content/merch/index.html)

## [DaCy v2.7.2](https://centre-for-humanities-computing.github.io/DaCy/)

[State-of-the-Art Danish NLP pipelines for spaCy](https://centre-for-humanities-computing.github.io/DaCy/)

[**🛸 spacy-transformers v1.3.1Sep 26, 2023** \\
Support for newer versions of Transformers](https://github.com/explosion/spacy-transformers/releases/tag/v1.3.1)

## [spaCy: a customizable NLP toolkit designed for developers](https://speakerdeck.com/sofievl/2023-06-15-odsc) [ODSC Europe](https://speakerdeck.com/sofievl/2023-06-15-odsc)

[**🦙 spacy-llm v0.1.0May 11, 2023** \\
Integrating LLMs into structured NLP pipelines](https://github.com/explosion/spacy-llm/releases/tag/v0.1.0)

## [Efficient Information Extraction From Text With spaCy](https://www.youtube.com/watch?v=1S8icpu9dX0) [JetBrains PyCharm](https://www.youtube.com/watch?v=1S8icpu9dX0)

[This webinar takes you through building a spaCy project that uses a named entity recognition (NER) model to extract entities of interest from restaurant reviews, like prices, opening hours and ratings.](https://www.youtube.com/watch?v=1S8icpu9dX0)

## [Predicting relations between SOAP note sections: The value of incorporating a clinical information model](https://www.sciencedirect.com/science/article/abs/pii/S1532046423000813?casa_token=xwJ-nM7yrPMAAAAA:lQmA8sCmWcHhxq9-0ducDxtT0lmsHVT185-7PjRAPTp-rXkbx5cx05KnzvJodXubWBl3Jhl5VgA) [Socrates, Gilson, Lopez, Chi, Taylor, Chartash (2023), Journal of Biomedical Informatics](https://www.sciencedirect.com/science/article/abs/pii/S1532046423000813?casa_token=xwJ-nM7yrPMAAAAA:lQmA8sCmWcHhxq9-0ducDxtT0lmsHVT185-7PjRAPTp-rXkbx5cx05KnzvJodXubWBl3Jhl5VgA)

[To support human annotation, we first annotate 100 Assessment and Plan subsections manually using Prodigy, and then use spacy-transformers to fine-tune a general domain RoBERTa-base model pretrained on OntoNotes 5 for both the Assessment and Plan section NER tagging.](https://www.sciencedirect.com/science/article/abs/pii/S1532046423000813?casa_token=xwJ-nM7yrPMAAAAA:lQmA8sCmWcHhxq9-0ducDxtT0lmsHVT185-7PjRAPTp-rXkbx5cx05KnzvJodXubWBl3Jhl5VgA)

## [Rulers, NER, and data iteration](https://blog.victoriaslocum.com/post/spanruler-ner-data)

[About the power of Rules + ML and the importance of iteration on your pipeline and your data.](https://blog.victoriaslocum.com/post/spanruler-ner-data)

## [Robust solutions with Explosion’s applied NLP philosophy](https://drive.google.com/file/d/1Mo5qEkTXxkg3DZ0tZZwBBFZ8t1ddymGB/view?usp=sharing) [UNC Charlotte](https://drive.google.com/file/d/1Mo5qEkTXxkg3DZ0tZZwBBFZ8t1ddymGB/view?usp=sharing)

## [Multi hash embeddings in spaCy](https://arxiv.org/abs/2212.09255) [Miranda, Kádár, Boyd, Van Landeghem, Søgaard, Honnibal (2022)](https://arxiv.org/abs/2212.09255)

[In this technical report we lay out a bit of history and introduce the embedding methods in spaCy in detail. Second, we critically evaluate the hash embedding architecture with multi-embeddings on Named Entity Recognition datasets from a variety of domains and languages. The experiments validate most key design choices behind spaCy’s embedders, but we also uncover a few surprising results.](https://arxiv.org/abs/2212.09255)

## [spaCy Cheat Sheet](https://github.com/explosion/assets/blob/main/spaCy/spaCy-cheat-sheet.pdf)

[Everything you need to know about spaCy as a handy two-page PDF.](https://github.com/explosion/assets/blob/main/spaCy/spaCy-cheat-sheet.pdf)

## [Introducing Holmes 4.0](/content/blog/introduction-to-holmes/index.html)

[A few weeks ago we released version 4.0 of Holmes, which we are now able to offer under a permissive MIT license. Holmes is a library in the spaCy Universe that runs on top of spaCy and enables information extraction and intelligent search, currently for English and German. Holmes goes beyond simple matching algorithms and allows you to look for a specified idea or ideas in a corpus of documents.](/content/blog/introduction-to-holmes/index.html)

## [Introducing spaCy v3.3](/content/blog/spacy-v3-3/index.html)

[spaCy v3.3 improves the speed of core pipeline components, adds a new trainable lemmatizer, and introduces trained pipelines for Finnish, Korean and Swedish.](/content/blog/spacy-v3-3/index.html)

## [Introducing spaCy Tailored Pipelines](/content/blog/introducing-spacy-tailored-pipelines/index.html)

[Explosion is pleased to announce a new development services offering, spaCy Tailored Pipelines. We’ll build you a custom natural language processing pipeline, delivered in a standardized format using spaCy’s projects system.](/content/blog/introducing-spacy-tailored-pipelines/index.html)

## [Neural edit-tree lemmatization for spaCy](/content/blog/edit-tree-lemmatizer/index.html)

[We are happy to introduce a new, experimental, machine learning-based lemmatizer\\
that posts accuracies above 95% for many languages. This lemmatizer learns to\\
predict lemmatization rules from a corpus of examples and removes the need to\\
write an exhaustive set of per-language lemmatization rules.](/content/blog/edit-tree-lemmatizer/index.html)

[**🌸 floret v0.10.0Oct 27, 2021** \\
fastText + Bloom embeddings for compact, full-coverage vectors with spaCy](https://github.com/explosion/floret)

## [Welcome spaCy to the Hugging Face Hub](https://huggingface.co/blog/spacy) [Hugging Face Blog](https://huggingface.co/blog/spacy)

[Hugging Face makes it really easy to share your spaCy pipelines with the community! With a single command, you can upload any pipeline package, with a pretty model card and all required metadata auto-generated for you.](https://huggingface.co/blog/spacy)

## [How We Found Pricey Provisions in New Jersey Police Contracts](https://www.propublica.org/article/how-we-found-pricey-provisions-in-new-jersey-police-contracts) [ProPublica](https://www.propublica.org/article/how-we-found-pricey-provisions-in-new-jersey-police-contracts)

[ProPublica and the Asbury Park Press scoured hundreds of police union agreements for details on publicly funded payouts to cops, using spaCy under the hood.](https://www.propublica.org/article/how-we-found-pricey-provisions-in-new-jersey-police-contracts)

## [Explosion in 2020: Our Year in Review](/content/blog/year-in-review-2020/index.html)

[While 2020 hasn’t been easy for anyone, at Explosion we’ve considered ourselves relatively fortunate in this most interesting year. We’ve always worked remotely, so we’ve been able to take both pride and comfort in continuing to ship good software. Here’s a look back at what we’ve been up to.](/content/blog/year-in-review-2020/index.html)

[**👑 spacy-streamlit v0.0.2Jun 23, 2020** \\
spaCy building blocks and visualizers for Streamlit apps](https://github.com/explosion/spacy-streamlit)

## [Training a custom entity linking model with spaCy](https://www.youtube.com/watch?v=8u57WSXVpmw)

[In this video, we show you how to create a custom Entity Linking model in spaCy to disambiguate different mentions of the person “Emerson” to unique identifiers in a knowledge base.](https://www.youtube.com/watch?v=8u57WSXVpmw)

## [Using spaCy with Hugging Face Transformers](https://speakerdeck.com/honnibal/spacy-meets-transformers) [PyCon India](https://speakerdeck.com/honnibal/spacy-meets-transformers)

[Transformer models like BERT have set a new standard for accuracy on almost every NLP leaderboard. However, these models are very new, and most of the software ecosystem surrounding them is oriented towards the many opportunities for further research. In this talk, Matt describes how you can now use these models in spaCy to work on real problems and the many opportunities transfer learningfor production NLP, regardless of which software packages you choose.](https://speakerdeck.com/honnibal/spacy-meets-transformers)

## [Intro to NLP with spaCy (2): Detecting programming languages](https://www.youtube.com/watch?v=KL4-Mpgbahw&list=PLBmcuObd5An559HbDr_alBnwVsGq-7uTF&index=2)

## [spaCy and Explosion: past, present & future](https://speakerdeck.com/inesmontani/spacy-and-explosion-past-present-and-future) [spaCy IRL 2019](https://speakerdeck.com/inesmontani/spacy-and-explosion-past-present-and-future)

## [Applied NLP: Lessons from the Field](https://docs.google.com/presentation/d/10wsqCTs4GqzWJlyrH2vhiwKSSPdmoP1N7D2Qr6MMKFk/edit\#slide=id.p) [spaCy IRL 2019](https://docs.google.com/presentation/d/10wsqCTs4GqzWJlyrH2vhiwKSSPdmoP1N7D2Qr6MMKFk/edit\#slide=id.p)

## [Introducing spaCy v2.1](/content/blog/spacy-v2-1/index.html)

[Version 2.1 of the spaCy Natural Language Processing library includes a huge number of features, improvements and bug fixes. In this post, we highlight some of the things we're especially pleased with, and explain some of the most challenging parts of preparing this big release.](/content/blog/spacy-v2-1/index.html)

## [Building new NLP solutions with spaCy and Prodigy](https://www.youtube.com/watch?v=jpWqz85F_4Y) [PyData Berlin](https://www.youtube.com/watch?v=jpWqz85F_4Y)

[“Commercial machine learning projects are currently like start-ups: many projects fail, but some are extremely successful, justifying the total investment. While some people will tell you to embrace failure, I say failure sucks — so what can we do to fight it? In this talk, I will discuss how to address some of the most likely causes of failure for new NLP projects.”](https://www.youtube.com/watch?v=jpWqz85F_4Y)

## [spaCy’s entity recognition model: incremental parsing with Bloom embeddings & residual CNNs](https://www.youtube.com/watch?v=sqDHBH9IjRU)

[spaCy v2.0’s Named Entity Recognition system features a sophisticated word embedding strategy using subword features and "Bloom" embeddings, a deep convolutional neural network with residual connections, and a novel transition-based approach to named entity parsing.](https://www.youtube.com/watch?v=sqDHBH9IjRU)

## [spaCy v1.0: Deep Learning with custom pipelines and Keras](/content/blog/spacy-deep-learning-keras/index.html)

[I'm pleased to announce the 1.0 release of spaCy, the fastest NLP library in the world. By far the best part of the 1.0 release is a new system for integrating custom models into spaCy. This post introduces you to the changes, and shows you how to use the new custom pipeline functionality to add a Keras-powered LSTM sentiment analysis model into a spaCy pipeline.](/content/blog/spacy-deep-learning-keras/index.html)

## [Introducing spaCy](/content/blog/introducing-spacy/index.html)

[Computers don't understand text. This is unfortunate, because that's what the web almost entirely consists of. We want to recommend people text based on other text they liked. We want to shorten text to display it on a mobile screen. We want to aggregate it, link it, filter it, categorise it, generate it and correct it. spaCy provides a library of utility functions that help programmers build such products.](/content/blog/introducing-spacy/index.html)

## [How Love Without Sound helps the music industry recover millions in revenue for artists with NLP, spaCy and Prodigy](/content/blog/love-without-sound-nlp-music-industry/index.html)

[A case study on Love Without Sound’s innovative AI-powered tools for the music industry and law firms specializing in royalty negotiations.](/content/blog/love-without-sound-nlp-music-industry/index.html)

## [Streaming spaCy](https://www.youtube.com/playlist?list=PLBmcuObd5An5_iAxNYLJa_xWmNzsYce8c)

[Join spaCy author and core developer Matt as he works on the library, develops features and fixes bugs, while chatting about all things NLP and open source. Every Thursday at 2pm CET and Friday at 11am CET.](https://www.youtube.com/playlist?list=PLBmcuObd5An5_iAxNYLJa_xWmNzsYce8c)

## [Applied NLP with LLMs: Beyond Black-Box Monoliths](https://speakerdeck.com/inesmontani/applied-nlp-with-llms-beyond-black-box-monoliths) [PyBerlin](https://speakerdeck.com/inesmontani/applied-nlp-with-llms-beyond-black-box-monoliths)

[In this talk, Ines shows some practical solutions for using the latest state-of-the-art models in real-world applications and distilling their knowledge into smaller and faster components.](https://speakerdeck.com/inesmontani/applied-nlp-with-llms-beyond-black-box-monoliths)

## [10 Years of Open Source: Navigating the Next AI Revolution](https://speakerdeck.com/inesmontani/10-years-of-open-source-navigating-the-next-ai-revolution) [EuroSciPy Keynote](https://speakerdeck.com/inesmontani/10-years-of-open-source-navigating-the-next-ai-revolution)

[In this talk, Ines shares the most important lessons we’ve learned in 10 years of working on open-source software, our core philosophies that helped us adapt to an ever-changing AI landscape and why open source and interoperability still wins over black-box, proprietary APIs.](https://speakerdeck.com/inesmontani/10-years-of-open-source-navigating-the-next-ai-revolution)

## [Toward Automatic Summarization of Hospital Discharge Notes](https://indigo.uic.edu/articles/thesis/Toward_Automatic_Summarization_of_Hospital_Discharge_Notes/27153255/1) [Landes (2024)](https://indigo.uic.edu/articles/thesis/Toward_Automatic_Summarization_of_Hospital_Discharge_Notes/27153255/1)

[For NLP tasks, vectorizers include spaCy token features such as part of speech (POS) tags, named entity recognition (NER) tags, dependency head relations and depth.](https://indigo.uic.edu/articles/thesis/Toward_Automatic_Summarization_of_Hospital_Discharge_Notes/27153255/1)

## [A practical guide to human-in-the-loop distillation](/content/blog/human-in-the-loop-distillation/index.html)

[This blog post presents practical solutions for using the latest state-of-the-art models in real-world applications and distilling their knowledge into smaller and faster components that you can run and maintain in-house.](/content/blog/human-in-the-loop-distillation/index.html)

## [Towards Structured Data: LLMs from Prototype to Production](https://speakerdeck.com/inesmontani/towards-structured-data-llms-from-prototype-to-production) [U.S. Census Bureau: Center for Optimization and Data Science Seminar](https://speakerdeck.com/inesmontani/towards-structured-data-llms-from-prototype-to-production)

[This talk presents pragmatic and practical approaches for how to use LLMs beyond just chat bots, how to ship more successful NLP projects from prototype to production and how to use the latest state-of-the-art models in real-world applications.](https://speakerdeck.com/inesmontani/towards-structured-data-llms-from-prototype-to-production)

[**🔌 prodigy-evaluate v0.1.0Mar 26, 2024** \\
Evaluate spaCy pipelines, print confusion matrices and more](https://github.com/explosion/prodigy-evaluate/releases/tag/v0.1.0)

## [Zero-Shot NER with GliNER and spaCy](https://www.youtube.com/watch?v=kPOtaXk-K-0) [Python Tutorials for Digital Humanities](https://www.youtube.com/watch?v=kPOtaXk-K-0)

[Tutorial by WJB Mattingly on how to integrate the generalist GLiNER model for Named Entity Recognition with spaCy's versatile NLP environment.](https://www.youtube.com/watch?v=kPOtaXk-K-0)

## [KAZU v1.5](https://github.com/AstraZeneca/KAZU)

[A biomedical NLP framework designed to handle production workloads, built by AstraZeneca and Korea University and using spaCy under the hood.](https://github.com/AstraZeneca/KAZU)

## [DeepZensols: A Deep Learning Natural Language Processing Framework for Experimentation and Reproducibility](https://aclanthology.org/2023.nlposs-1.16/) [Landes, Di Eugenio, Caragea (2023)](https://aclanthology.org/2023.nlposs-1.16/)

[A linguistic feature mapper that translates spaCy to wordpieces, which are token sub-units with associated vectors, is also accessible as an easy to configure module.](https://aclanthology.org/2023.nlposs-1.16/)

## [calamanCy: A Tagalog Natural Language Processing Toolkit](https://arxiv.org/abs/2311.07171) [Miranda (2023), EMNLP 2023](https://arxiv.org/abs/2311.07171)

[We introduce calamanCy, an open-source toolkit for constructing NLP pipelines for Tagalog. It is built on top of spaCy, enabling easy experimentation and integration with other frameworks.](https://arxiv.org/abs/2311.07171)

## [scispacy v0.5.3](https://allenai.github.io/scispacy/)

[A Python package containing spaCy models for processing biomedical, scientific or clinical text, developed by AI2.](https://allenai.github.io/scispacy/)

[**🦙 spacy-llm v0.5.0Sep 8, 2023** \\
Improved user API and novel Chain-of-Thought prompting for more accurate NER](https://github.com/explosion/spacy-llm/releases/tag/v0.5.0)

## [How Good is the Model in Model-in-the-loop Event Coreference Resolution Annotation?](https://arxiv.org/abs/2306.05434) [Ahmed, Nath, Regan, Pollins, Krishnaswamy, Martin (2023)](https://arxiv.org/abs/2306.05434)

[Figure 6 illustrates the interface design of the annotation methodology on the popular model-in-the-loop annotation tool - Prodigy. We use this tool for the simplicity it offers in plugging in the various ranking methods we explained.](https://arxiv.org/abs/2306.05434)

## [spaCy Plugin for VSCode](https://marketplace.visualstudio.com/items?itemName=Explosion.spacy-extension)

[The spaCy VSCode Extension provides additional tooling and features for working with spaCy’s config files. Version 1.0.0 includes hover descriptions for registry functions, variables, and section names within the config as an installable extension.](https://marketplace.visualstudio.com/items?itemName=Explosion.spacy-extension)

## [Intro to NLP with spaCy for Digital Humanities](https://speakerdeck.com/victorialslocum/princeton-workshop-presentation) [Princeton University](https://speakerdeck.com/victorialslocum/princeton-workshop-presentation)

## [The Nesta Skills Extractor Library](https://www.escoe.ac.uk/the-skills-extractor-library/) [Economic Statistics Centre of Excellence](https://www.escoe.ac.uk/the-skills-extractor-library/)

[A new library for extracting skills from job adverts and mapping them to a taxonomy of your choice, built on top of spaCy.](https://www.escoe.ac.uk/the-skills-extractor-library/)

## [Training spaCy NER Models with Prodigy](https://github.com/explosion/assets/blob/main/Prodigy/Prodigy_NER_flowchart_v2_0_0_light.pdf)

[This handy flowchart contains our most common tips, tricks, and best practices for training and updating spaCy named entity recognition models with Prodigy.](https://github.com/explosion/assets/blob/main/Prodigy/Prodigy_NER_flowchart_v2_0_0_light.pdf)

[**🛸 spacy-transformers v1.2.0Jan 14, 2023** \\
Better alignment for fast tokenizers](https://github.com/explosion/spacy-transformers/releases/tag/v1.2.0)

## [The triangulation of ethical leader signals using qualitative, experimental, and data science methods](https://www.sciencedirect.com/science/article/abs/pii/S1048984322000613) [Banks, Ross, Toth, Tonidandel, Goloujeh, Dou, Wesslen (2022)](https://www.sciencedirect.com/science/article/abs/pii/S1048984322000613)

[This additional text was labeled by the same coding team using Prodigy, \[...\] a flexible user interface tool built on top of spaCy, a leading open source library in python for natural language processing. We created a spaCy end‐to‐end project workflow including package versioning, data pre‐processing, data ingestion into a database, annotation sessions using Prodigy’s user interface, model training, model evaluation, python packaging, and visual app for testing the model.](https://www.sciencedirect.com/science/article/abs/pii/S1048984322000613)

## [How the Guardian approaches quote extraction with NLP](/content/blog/guardian/index.html)

[A case study of the Guardian's spaCy-Prodigy workflow to modularize quote extraction for content creation. This study includes iterative annotation guidelines and custom interface functionality.](/content/blog/guardian/index.html)

## [Introducing spaCy v3.4](/content/blog/spacy-v3-4/index.html)

[spaCy v3.4 brings typing and speed improvements along with new vectors for English CNN pipelines and new trained pipelines for Croatian.](/content/blog/spacy-v3-4/index.html)

## [How we built a Stack Overflow Community questions analyzer](https://about.gitlab.com/blog/2022/04/28/how-we-built-a-stack-overflow-community-questions-analyzer-and-you-can-too/) [GitLab Blog](https://about.gitlab.com/blog/2022/04/28/how-we-built-a-stack-overflow-community-questions-analyzer-and-you-can-too/)

[How GitLab used spaCy to analyze and better understand Stack Overflow community questions about their tools and products.](https://about.gitlab.com/blog/2022/04/28/how-we-built-a-stack-overflow-community-questions-analyzer-and-you-can-too/)

## [When Women Make Headlines](https://pudding.cool/2022/02/women-in-headlines/) [The Pudding](https://pudding.cool/2022/02/women-in-headlines/)

[Using spaCy and other packages from the NLP ecosystem for analyzing more than 382,000 headlines to see how women are represented (or misrepresented) in the news.](https://pudding.cool/2022/02/women-in-headlines/)

## [Talking sense: using machine learning to understand quotes](https://www.theguardian.com/info/2021/nov/25/talking-sense-using-machine-learning-to-understand-quotes) [The Guardian Blog](https://www.theguardian.com/info/2021/nov/25/talking-sense-using-machine-learning-to-understand-quotes)

[How the Guardian uses spaCy and Prodigy to train a machine learning model that helps extract quotes from news articles and match them to the correct source.](https://www.theguardian.com/info/2021/nov/25/talking-sense-using-machine-learning-to-understand-quotes)

[**🛸 spacy-transformers v1.1.0Oct 18, 2021** \\
Better serialization, full ModelOutput, mixed-precision training and more](https://github.com/explosion/spacy-transformers/releases/tag/v1.1.0)

## [Introduction to Japanese Natural Language Processing](https://www.japanesenlp.com/) [Masato Hagiwara, Paul O’Leary McCann (2021)](https://www.japanesenlp.com/)

[A thorough guide for programmers working with Japanese text, covering fundamental issues like tokenization and recent research topics like generating natural language texts.](https://www.japanesenlp.com/)

## [spaCy v3: Custom trainable relation extraction component](https://www.youtube.com/watch?v=8HL-Ap5_Axo)

[spaCy v3.0 features new transformer-based pipelines that get spaCy’s accuracy right up to the current state-of-the-art, and a new training config and workflow system to help you take projects from prototype to production. In this video, Sofie shows you how to apply all these new features when implementing a custom trainable component from scratch.](https://www.youtube.com/watch?v=8HL-Ap5_Axo)

## [The Physical Traits that Define Men and Women in Literature](https://pudding.cool/2020/07/gendered-descriptions/) [The Pudding](https://pudding.cool/2020/07/gendered-descriptions/)

[Analysis of physical traits most tied to gender in literature using spaCy.](https://pudding.cool/2020/07/gendered-descriptions/)

[**🛸 spacy-transformers v0.6.0May 24, 2020** \\
Update to transformers v2.5.0](https://github.com/explosion/spacy-transformers/releases/tag/v0.6.0)

## [Intro to NLP with spaCy (4): Detecting programming languages](https://www.youtube.com/watch?v=IqOJU1-_Fi0&list=PLBmcuObd5An559HbDr_alBnwVsGq-7uTF&index=4)

## [spaCy and the future of multi-lingual NLP](https://speakerdeck.com/inesmontani/spacy-and-the-future-of-multi-lingual-nlp) [META Forum](https://speakerdeck.com/inesmontani/spacy-and-the-future-of-multi-lingual-nlp)

## [Millennials Kill Everything](https://pudding.cool/2019/09/millennials/) [The Pudding](https://pudding.cool/2019/09/millennials/)

[Analysis on media reporting of millenials using spaCy. From napkins to marriage to Applebees, just looking at headlines you’d guess that for the past decade the millennial generation’s been on a rampage.](https://pudding.cool/2019/09/millennials/)

## [David Dodson: spaCy in the News: Quartz’s NLP pipeline](https://www.youtube.com/watch?v=azrVX8JksMU&list=PLBmcuObd5An4UC6jvK_-eSl6jCvP1gwXc&index=11) [spaCy IRL 2019](https://www.youtube.com/watch?v=azrVX8JksMU&list=PLBmcuObd5An4UC6jvK_-eSl6jCvP1gwXc&index=11)

## [Entity linking functionality in spaCy](https://drive.google.com/file/d/1EuGxcQLcXvjjkZ-KRUlwpr_doBVyEBEG/view) [spaCy IRL 2019](https://drive.google.com/file/d/1EuGxcQLcXvjjkZ-KRUlwpr_doBVyEBEG/view)

## [FAQ \#1: Tips & tricks for NLP, annotation & training with Prodigy and spaCy](https://www.youtube.com/watch?v=tMAU3gLbKII)

[In this video, Ines talks about a few frequently asked questions and shares some general tips and tricks for how to structure your NLP annotation projects, how to design your label schemes and how to solve common problems.](https://www.youtube.com/watch?v=tMAU3gLbKII)

## [Can You Verifi This? Studying Uncertainty and Decision-Making About Misinformation](https://wesslen.netlify.app/publication/icwsm-2018-1/) [Karduni, Wesslen, Santhanam, Cho, Volkova, Arendt, Shaikh, Dou (2018)](https://wesslen.netlify.app/publication/icwsm-2018-1/)

[HCI interface to identify misinformation on social media using spaCy for NER.](https://wesslen.netlify.app/publication/icwsm-2018-1/)

## [Introducing custom pipelines and extensions for spaCy v2.0](/content/blog/spacy-v2-pipelines-extensions/index.html)

[As the release candidate for spaCy v2.0 gets closer, we've been excited to implement some of the last outstanding features. One of the best improvements is a new system for adding pipeline components and registering extensions to the Doc, Span and Token objects. In this post, we'll introduce you to the new functionality, and finish with an example extension package, spacymoji.](/content/blog/spacy-v2-pipelines-extensions/index.html)

## [Multi-threading spaCy's parser and named entity recognizer](/content/blog/multithreading-with-cython/index.html)

[In v0.100.3, we quietly rolled out support for GIL-free multi-threading for spaCy's syntactic dependency parsing and named entity recognition models. Because these models take up a lot of memory, we've wanted to release the global interpretter lock (GIL) around them for a long time. When we finally did, it seemed a little too good to be true, so we delayed celebration — and then quickly moved on to other things. It's now past time for a write-up.](/content/blog/multithreading-with-cython/index.html)

### Newsletter

Join our mailing list to receive updates about new blog posts and projects!

💌

[View past issues](https://us12.campaign-archive.com/home/?u=83b0498b1e7fa3c91ce68c3f1&id=ecc82e0493)
