Lilac

Dataset curation for LLM training

Code & Dev Tools Open Source Open Source
Researched · Published
RECATOOLS Score
6.5 / 10
Capability
7
Value for money
7
Ease of use
6
ASEAN readiness
5
API quality
6
Founded
2023
HQ
San Francisco, California, USA
Users
Launched
Developer

Overview

Lilac (acquired by Databricks in 2024) is an open-source tool for inspecting, curating and labeling large datasets used to train and fine-tune LLMs. Particularly useful for finding PII, duplicates and quality issues at scale.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 20 May 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Free
Free
Free tier with core features.

Use cases

Dataset inspection Training-data curation PII detection

What you can produce with Lilac

  • Cluster a large text dataset and get concise auto-generated titles for each cluster to see its major segments.
  • Scan a training or fine-tuning corpus for PII such as emails, phone numbers, and secrets before releasing or training on it.
  • Detect near-duplicate documents in a dataset and filter them out to reduce training redundancy.
  • Run semantic and keyword search over millions of documents through a local browser UI.
  • Annotate and tag data points with custom concepts, then export the filtered subset for fine-tuning.
  • Compute text statistics and quality signals across a dataset via the Python API to drive programmatic curation.
Advertisement

ASEAN Perspective

Lilac in Southeast Asia

ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).

RECATOOLS Verdict

Lilac is an open-source tool for exploring, clustering, searching and cleaning unstructured text datasets — useful for LLM evaluation and for preparing data for RAG, fine-tuning and pre-training. Built by ex-Google engineers, it suits data and ML teams who need to understand and curate large text corpora rather than treat them as a black box.

The critical caveat is status: Lilac was acquired by Databricks, and its capabilities are being folded into the Databricks platform, so the standalone open-source project is effectively in maintenance/archive mode. Evaluate it as either a Databricks feature or an OSS reference rather than an actively developed independent product. Self-hostable for free; English-only; no SEA-specific provisions.

Independent AI-assisted assessment by RECATOOLS.

What people say

Lilac's story as a standalone product is over: Databricks acquired the Boston-based startup in March 2024 to fold its dataset-curation technology into the Mosaic AI stack, and the open-source repository (now under the databricks GitHub org) was archived and made read-only on 25 July 2025. The pip package still installs and the documentation remains online, but there is no active maintenance, no roadmap, and the hosted Lilac Garden service was absorbed into Databricks. Anyone evaluating it in 2026 should treat it as legacy software.

That ending should not obscure how well-regarded the tool was. In its roughly two years of independent life, Lilac earned genuine respect among LLM practitioners for making unstructured text datasets explorable: it could cluster and auto-title a million data points in about twenty minutes on its accelerated Garden service (the team claimed 100x over local computation), surface PII, near-duplicates, profanity, and quality issues at scale, and support semantic and concept-based search over training corpora. It was used to dissect well-known public datasets — OpenOrca, UltraChat, LMSYS Chatbot Arena logs — and its cluster visualizations circulated widely in the fine-tuning community. Users liked that it ran on-device with a local UI and Python API, keeping sensitive data off third-party servers.

The complaints were the usual ones for a young open-source tool — rough edges at very large scale, compute-hungry embedding steps without Garden — but sentiment was broadly positive, which is partly why Databricks bought it.

Today Lilac fits almost no one as a new adoption: Databricks customers get its DNA through Mosaic AI's data tooling, while teams wanting an actively maintained open alternative for dataset exploration and curation have moved to tools like Nomic Atlas, Argilla, or Cleanlab. The archived code remains useful mainly for reference or one-off local analysis.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on

This entry was compiled from publicly available data including Lilac's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Lilac unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Lilac directly →

Spotted something out of date? Suggest an update →

Advertisement