Community call: OpenRefine LLM Extension Demo

We are organizing an OpenRefine community call focused on concrete uses of the OpenRefine LLM Extension.

This session is intended for anyone interested in using LLMs with OpenRefine. The goal is to show practical examples of how the extension can support data-wrangling workflows, including use cases from data journalism and library metadata work.

This session will be recorded and shared on YouTube.

Session Overview

The OpenRefine LLM Extension makes it possible to call large language models in OpenRefine workflows. The session will show two examples of how LLMs are used in journalism and GLAM use cases, including:

  • Preparing and transforming data for journalistic work
  • Exploring library-specific metadata workflows
  • Using local or external providers, including Ollama and other LLM services
  • Identifying what works well and what still needs improvement

Questions from users, trainers, extension developers, and contributors are welcome.

Agenda

  • Short introduction of each participant
  • Community announcements
  • Demo by @h_piedcoq based on his DataHarvest 2026 tutorial, with a focus on data journalism workflows
  • Demo by @psm on the use of OpenRefine with the AI plugin for library-specific purposes, including integrations with Ollama and other providers
  • Questions and discussion with @Sunil_Natraj, developer of the OpenRefine LLM Extension

The recording of the OpenRefine Community Call: OpenRefine LLM Extension Demo is now available on YouTube:

Thank you again to @h_piedcoq, @psm, and @Sunil_Natraj for contributing to the session.

Recording overview

The recording includes three main parts:

  • Hervé Piedcoq’s demo (starts at 00:19 min), based on his DataHarvest 2026 tutorial, with a focus on data journalism workflows. Hervé spent a significant part of the demo showing name cleaning with the LLM extension. His tutorial materials are linked in the first post of this thread.

  • Parthasarathi Mukhopadhyay’s demo (starts at 24:23 min), focused on library-specific use cases and testing different models via Ollama. His related materials are available here:

  • Community discussion (starts at 50:15 min), including questions with Sunil Natraj, developer of the OpenRefine LLM Extension.

Topics covered

Some of the points discussed during the session include:

  • using the LLM extension for name cleaning and data journalism workflows
  • testing local models through Ollama
  • library metadata use cases
  • using structured JSON responses so that an LLM can return multiple data points in a single answer
  • including fields such as confidence scores or reasoning notes in structured outputs
  • possible uses of LLMs to review reconciliation service candidates