Deterministic, interpretation-free job-posting corpus: ingested daily at 06:00 UTC from 38,000+ public machine-readable feeds (ATS and board structured endpoints), ~50k records/day in and ~25k net-new postings/day after dedup. The pipeline is explicitly non-inferential — translate to Schema.org JobPosting form, best-effort ISO normalisation applied ONLY where upstream provides unambiguous explicit structure, original upstream values preserved whenever a field is ambiguous or unmappable, RFC 8785 canonicalisation, SHA-256 byte-level deduplication, committed to a content-addressed store with stable IDs. No scraping, no LLM extraction, no enrichment, no classification, no inferred metadata, and no fabricated fields (if upstream omits a field, it is omitted). Served through a keyed REST search endpoint filterable on title, location, industry, company, employment_type, job_location_type (Onsite/Hybrid/Remote), language (ISO 639-2), country (ISO 3166), currency, salary_min/max and posted_after, with key-repetition OR / key-difference AND query semantics. The deliberate anti-thesis of AI-parsed job feeds: it sells the absence of an interpretation layer.
Your sourcing agent asks Kaleh and files the sample in your catalog — private to you.
Universe, instruments and categories.
Kaleh has not added their own details yet. Not yet on file: