Datasets
Buyer room
KALEH.NETManage supplier listing
Trace Jobs Core (Kaleh LLC) — Deterministic, interpretation-free job-posting corpus: ingested daily at 06:00 UTC from 38,000+ public machine-readable feeds (ATS and board structured endpoints), ~50k records/day in and ~25k net-new postings/day after dedup. The pipeline is explicitly non-inferential — translate to Schema.org JobPosting form, best-effort ISO normalisation applied ONLY where upstream provides unambiguous explicit structure, original upstream values preserved whenever a field is ambiguous or unmappable, RFC 8785 canonicalisation, SHA-256 byte-level deduplication, committed to a c…
supplier index read from their site
Deterministic, interpretation-free job-posting corpus: ingested daily at 06:00 UTC from 38,000+ public machine-readable feeds (ATS and board structured endpoints), ~50k records/day in and ~25k net-new postings/day after dedup. The pipeline is explicitly non-inferential — translate to Schema.org JobPosting form, best-effort ISO normalisation applied ONLY where upstream provides unambiguous explicit structure, original upstream values preserved whenever a field is ambiguous or unmappable, RFC 8785 canonicalisation, SHA-256 byte-level deduplication, committed to a content-addressed store with stable IDs. No scraping, no LLM extraction, no enrichment, no classification, no inferred metadata, and no fabricated fields (if upstream omits a field, it is omitted). Served through a keyed REST search endpoint filterable on title, location, industry, company, employment_type, job_location_type (Onsite/Hybrid/Remote), language (ISO 639-2), country (ISO 3166), currency, salary_min/max and posted_after, with key-repetition OR / key-difference AND query semantics. The deliberate anti-thesis of AI-parsed job feeds: it sells the absence of an interpretation layer.
Scale 38From 2 yearsCoverage Information Technology · Industrials
Sample, licence terms, pricing and eval results when Kaleh publishes them. Until then, discover alternatives today with a 7-day trial.
Start 7-day trial →What buyers ask about Kaleh, answered from this page.
Trace Jobs Core (Kaleh LLC) — Deterministic, interpretation-free job-posting corpus: ingested daily at 06:00 UTC from 38,000+ public machine-readable feeds (ATS and board structured endpoints), ~50k records/day in and ~25k net-new postings/day after dedup. The pipeline is explicitly non-inferential — translate to Schema.org JobPosting form, best-effort ISO normalisation applied ONLY where upstream provides unambiguous explicit structure, original upstream values preserved whenever a field is ambiguous or unmappable, RFC 8785 canonicalisation, SHA-256 byte-level deduplication, committed to a c…
Kaleh offers (Alternative, Reference, Sentiment) — 38,000+ structured feeds ingested daily at 06:00 UTC; ~50,000 records/day ingested, ~25,000 net-new postings/day after SHA-256 deduplication — implying roughly 7-9M new unique postings per year accumulating in an immutable content-addressed index described by its founder as 'millions of records'. Growth is measurable: 9,800 feeds and ~13k new postings/day at the Hacker News launch, versus 38,000+ feeds and ~25k/day now — approximately 4x feed coverage and ~2x net-new volume in the interval..
Vendor-named buyers: investment and market analysts tracking sectors/titles/regional hiring directly from Python/Pandas without unstructured text; AI/RAG product teams feeding structured fields into embedding models and vector DBs to eliminate token waste and HTML-derived hallucination; niche job boards and aggregators syncing net-new listings to Postgres/Supabase on cron in place of broken Puppeteer/Playwright scrapers. The published example set prebuilds multi-country trend matrices, remote-share lead generation, daily incremental sync, salary-threshold compensation analysis and top-paying-role ranking. Distinctive extended use: as a CORPUS HYGIENE BASELINE — because it is deterministically derived, content-hashed and inference-free, it is an unusually clean reference set for measuring the field-accuracy drift and hallucination rate of LLM-parsed job datasets.
The data is with 2 years of history.
Coverage spans US, APAC; Application Software, Human Capital & Employment Services; alternative, reference, sentiment; equities.
Kaleh has not added their own details yet. Not yet on file: