Indic-first embedding models, with provenance you can audit.
We train small, sharp retrieval models for Indian languages and domains — and we can tell you where every training example came from. Provenance is not a footnote here; it is the product.
| Model | What it is |
|---|---|
multilingual-embedding |
Flagship general-purpose multilingual retriever. One meaning searched across many languages. |
embed-legal-en |
English Supreme Court judgment retriever. +76% in-distribution Recall@1 (0.309→0.545, confidence intervals disjoint). Trained only on statutory public-domain judgment text; scope is English judgments, and the card says so. |
Domain specialists follow the same discipline — legal and finance retrievers tuned for Indian text — and ship when the model earns its scope, not before.
Don't take the numbers on faith — run them. The public playground lets you search a single meaning across languages and watch a domain adapter separate the right answer from the noise: