Item detail
github.com

kreuzberg-dev/kreuzberg

kreuzberg-dev/kreuzberg is a framework that RepoRadar is tracking in its Radar section, currently rated Gold tier with a 'try now' verdict. Its strongest signal is workflow potential, scored 10.0 out of 10.

Score8.9
Popularity82.0
Riskconditional
TierGold
Score breakdown
Usefulness9.0
Novelty8.0
Momentum8.0
Maturity8.6
Open-source/build7.4
Evidence7.2
Workflow potential10.0
Setup ease6.4

Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.

Why it matters

Useful for teams building document-heavy AI workflows that need one serious extraction layer instead of a pile of single-format parsers and ad hoc OCR scripts.

Where this stands now

kreuzberg-dev/kreuzberg ranks #9 of 1508 tracked Radar items by composite score (8.9 against a section median of 7.8). The section currently carries 643 Gold, 443 Silver, 422 Bronze. RepoRadar has retained observations for this record since 2026-06-19 (88 days in the current window).

Who should use it

RAG builders document automation teams MCP tool builders developers who need extraction across multiple languages

Who should skip it

Consider kreuzberg-dev/kreuzberg lower priority if you already have a working solution in this category.

About this signal

kreuzberg-dev/kreuzberg is tracked by RepoRadar as a framework in the Radar section. First seen 2026-06-19; the source record was last checked on 2026-06-19. The current verdict is 'try now' with a Gold tier and moderate setup difficulty. The standout signals for kreuzberg-dev/kreuzberg are workflow potential (10.0) and practical usefulness (9.0), while setup ease (6.4) trails — that balance shapes where it fits best. This page summarizes the public evidence on the linked source page and states where additional review is still needed.

How this item is evaluated

The kreuzberg-dev/kreuzberg record combines a 8.9/10 composite score with separate popularity (82.0), risk (conditional), and setup (moderate) signals. See the scoring methodology for the current weights and evidence definitions.

Questions worth asking before you adopt this

Putting this into practice? Read How to evaluate an AI tool before you adopt it for the checklist behind this score.

Risk explanation

The project can ingest private documents and optionally connect to hosted OCR or LLM providers, so verify both the data path and the Elastic License 2.0 usage terms before adopting it in a commercial workflow.

Evidence links
Closest alternatives / related signals
document-intelligence ocr pdf-extraction mcp rag
Verification record

What RepoRadar actually verified

Tested in a bounded workflow Source-only update since test

Bounded representative workflow retained by RepoRadar verification harness. Last checked 2026-07-13T10:37:30.605699Z.

A source release dated 2026-09-14 is newer than the retained test dated 2026-07-13; the previous test may no longer represent the current release.

Read as a sequence: 1 retained check(s) on 2026-07-13 with outcomes of 1 passed; evidence scopes covered bounded representative workflows. Recorded setup time totals 1 minute(s).

passed · cohort-20260712-kreuzberg-html-extraction-workflow

Tester
RepoRadar automated local verification harness
Started
2026-07-13T10:37:28.017634Z
Completed
2026-07-13T10:37:30.605699Z
Environment
Windows 10 AMD64; Python 3.11.9; credential-stripped child environment; disposable home/cache
Install/setup time
1 minute(s)
Evidence scope
Bounded representative workflow
Cleanup
Per-check temporary home and work directory removed. Shared cohort package cache removed.
Actions exercised
  • Created a disposable home, work directory, and isolated package cache with credential-like environment variables excluded.
  • Created 2 synthetic fixture file(s) inside the disposable work directory; retained hashes prove the exact inputs.
  • Created a local HTML runbook containing a title, two ordered recovery steps, styling, and a script-only exclusion marker.
  • Ran Kreuzberg synchronously with cache and OCR disabled, asserted meaningful text was retained and script content excluded, and exported the result.
  • Executed bounded check: Extract a synthetic HTML incident runbook into Markdown with Kreuzberg's Python API.
  • Captured the complete sanitized stdout, stderr, exit status, artifact checks, and 2.59-second wall time.
Observed results
  • Command exited 0 after 2.59 seconds.
  • Kreuzberg extracted the runbook title and both recovery steps while omitting script content from the retained Markdown result.
  • Expected marker 'CHECK_OK format=markdown title=true steps=2' was observed in retained output.
  • Validated result.json: 4 required marker(s) present and 1 excluded marker(s) absent; size and SHA-256 are retained.
Observed strengths
  • The Python API provided local HTML-to-Markdown extraction with explicit cache/OCR controls and clean content filtering.
Friction
  • The official Python wheel is large because it bundles native extraction support beyond the one HTML format exercised here.
  • Setup or runtime emitted 3 stderr line(s); the complete warnings/errors are preserved in the retained log.
Limitations
  • The HTML-only, OCR-disabled fixture validates local structured text extraction, not PDF/Office parsers, OCR, tables, chunking, language detection, or batch throughput.
  • This credential-free disposable workflow does not establish production scale, model quality, reliability under sustained use, or team adoption.

Pricing assessment: The local open-source extractor used no OCR service, hosted parser, account, or model.

Privacy assessment: The synthetic HTML was read only from the disposable work directory and its extracted text was never sent externally.

Open retained test log →

Verification sources

Longitudinal intelligence

How this decision record is moving

Raw history JSON →

69 dated snapshots retained from 2026-06-19 through 2026-09-15; see the snapshot index for explicit coverage gaps. Stars, version, release, pricing, integration, risk, maintenance, verdict, score, and momentum fields remain explicit even when a source has not reported them. Repository momentum is a normalized 0–10 RepoRadar signal; GitHub stars appear only where the popularity monitor retained exact timestamped observations.

RepoRadar score8.9 current · +0.0 net
Repository momentum9.3 current · +1.3 net
GitHub stars (observed)9,308 current · +666 net
GitHub stars9,308 exact observation
Versionv1.2.1
Last release2026-09-14T10:15:39Z
Maintenanceactive
Current riskconditional
Current verdicttry now
Pricing baselineNo structured commercial pricing baseline
Pricing checkedNot applicable or not recorded
Pricing freshnessNo dated commercial pricing review
Integrations baselineNo structured integrations recorded

Recent dated points

DateScoreMomentumStarsRiskVerdictMaintenance
2026-09-158.99.39,308conditionaltry nowactive
2026-09-138.99.09,296conditionaltry nowactive
2026-09-128.99.09,296conditionaltry nowactive
2026-09-118.99.39,291conditionaltry nowactive
2026-09-088.99.09,275conditionaltry nowactive
2026-09-078.98.0Not recordedconditionaltry nownot recorded
2026-09-058.99.39,264conditionaltry nowactive
2026-09-048.99.39,264conditionaltry nowactive
2026-09-038.98.0Not recordedconditionaltry nownot recorded
2026-09-028.99.09,247conditionaltry nowactive
2026-09-018.99.09,247conditionaltry nowactive
2026-08-308.99.39,234conditionaltry nowactive

Why the record changed

version changed

Version changed: v1.1.5 → v1.2.1.

stars changed

Stars changed: 9296 → 9308.

stars changed

Stars changed: 9294 → 9296.

stars changed

Stars changed: 9291 → 9294.

version changed

Version changed: v1.1.2 → v1.1.5.

stars changed

Stars changed: 9275 → 9291.

stars changed

Stars changed: 9234 → 9247.

stars changed

Stars changed: 9221 → 9234.

stars changed

Stars changed: 9219 → 9221.

stars changed

Stars changed: 9218 → 9219.

stars changed

Stars changed: 9215 → 9218.

stars changed

Stars changed: 9216 → 9215.