Skip to main content
Conare’s embedding and rerank models also work without namespaces, through OpenAI- and Cohere-compatible endpoints. Point your existing client at Conare. A service that only embeds and reranks can use a Models only key, which reaches no namespace.

Embed

POST /v1/embeddings
Response

Request

  • input (string or array, required): 1-128 texts.
  • model (string, default e5-mine): e5-mine for text you store, e5-mine-query for searches.
  • encoding_format (string, default float): float or base64.

Response

  • data: One embedding per input, in order. Each has 1,024 numbers (shortened above).
  • usage: The tokens read.
Only the first 6,000 characters of each text are embedded. A query’s vector includes today’s date, so it changes from day to day. LangChain’s OpenAIEmbeddings needs check_embedding_ctx_length=False, or it sends token ids instead of text.

Rerank

POST /v1/rerank
Response

Request

  • query (string, required): Up to 8,192 characters.
  • documents (array, required): 1-100 documents, as strings or { "text": ... } objects. Each is up to 24,000 characters, and together with the query up to 256,000.
  • top_n (integer, default all): Return only the best top_n.
  • return_documents (boolean, default false): Add each document’s text to its result.

Response

  • results: Best first. index is the document’s position in documents, and relevance_score is higher for a better match.
  • unscored: Documents the reranker didn’t reach in time. They aren’t in results.
  • usage: The tokens read.
Use Cohere’s v1 client, cohere.Client. cohere.ClientV2 calls /v2/rerank, which Conare doesn’t serve.

Models

GET /v1/models lists the models your key can call. Models trained for your organization are listed too.

With the Conare SDK

These have the same limits as the /v1 routes.