An HTTP 200 from Ollama does not prove that an embedding represents the whole chunk you stored. The /api/embed contract defaults to truncate: true: oversized inputs can be cut to fit the context window. If your index saves the original text beside that vector, retrieval searches a representation of a shorter passage while your application displays the full one.
For a document-ingestion pipeline, prefer an error. Send truncate: false, quarantine rejected chunks, and split them before inserting anything. A missing vector is visible. A vector that quietly ignores the paragraph containing the answer is much harder to diagnose.

This is a documented failure risk, not a claim that every Ollama release silently truncates every oversized request. A GitHub report describes versions that returned overflow errors even with truncation enabled. That distinction matters: test the installed server, and build an ingestion path that can handle rejection either way.
Make oversize chunks fail before insertion
Begin with the current endpoint, POST /api/embed. It accepts input as a string or an array and returns embeddings as an array of vectors. Old examples using /api/embeddings, prompt, and a singular embedding describe a different request and response shape. Do not change the URL while leaving the old parser in place.
For Nomic v1.5, the model owner also requires task prefixes. Index documents with search_document: and embed questions with search_query:. Those prefixes belong to this model's usage instructions; they are not universal Ollama syntax.
A short request establishes the contract:
ollama pull nomic-embed-text:v1.5
curl --fail-with-body http://localhost:11434/api/embed \
-H 'Content-Type: application/json' \
-d '{"model":"nomic-embed-text:v1.5","input":["search_document: The staging credential expires on Friday."],"truncate":false}'
Require a successful response and one vector for that one input before inserting the record. An HTTP error belongs in a rejected-work queue with the document ID and chunk location. Keep the full text in your source store, but do not create a searchable row that pretends the failed chunk has a vector.
The Python boundary can stay small:
import math
import ollama
MODEL = "nomic-embed-text:v1.5"
def embed_checked(texts, expected_dim):
result = ollama.embed(
model=MODEL,
input=["search_document: " + text for text in texts],
truncate=False,
)
vectors = result["embeddings"]
if len(vectors) != len(texts):
raise ValueError("embedding count mismatch")
for vector in vectors:
if len(vector) != expected_dim:
raise ValueError("embedding dimension mismatch")
if not all(math.isfinite(x) for x in vector):
raise ValueError("non-finite embedding")
if not any(x != 0 for x in vector):
raise ValueError("zero embedding")
return vectors
Supply expected_dim from the schema of the collection you are writing. Establish it when you create that collection using a known short input and the selected model configuration. Do not silently accept whatever dimension a later response happens to return. A model or dimension change needs its own index migration.
This guard checks the response shape and rejects obviously unusable vectors. It cannot prove semantic quality. It also deliberately lets API exceptions reach the caller: catch them at the job boundary, record the failure, and leave the vector collection unchanged. Do not turn an exception into an empty vector to keep the import counter moving.
For larger batches, validate the complete returned batch before any insertion. If an oversized item makes the request fail, isolate rejected items in smaller requests, then split the offending text. Reduce batch size when diagnosing memory pressure or malformed inputs. Never assume a partially completed batch can be paired with all original chunks by position.
The model card is not the serving contract
Nomic's v1.5 model card lists a sequence length of 8192. The Ollama library currently labels its v1.5 artifact with a 2K context window. These are different statements about different layers. Neither a leaderboard nor a model-card maximum tells you what a particular server will accept today.
Inspect the deployed tag with ollama show nomic-embed-text:v1.5, record the Ollama version, and probe the actual API. Avoid copying an 8192-token claim into your splitter configuration without checking the serving path. A context override that works in one backend is not evidence that another backend uses the same allocation or scaling behavior.
Chunk by the tokenizer used by your embedding stack when it exposes one. Word counts and character counts are useful starting heuristics, not token guarantees. Code, long identifiers, multilingual text, and task prefixes change the budget. Leave room for the prefix and model-added tokens, then let truncate: false enforce the server boundary.
A useful acceptance test has a short control input and an intentionally oversized input. The control should yield a finite, nonzero vector with the collection's expected dimension. The oversized case must either fail before insertion or be split into accepted pieces by your application. Record the actual behavior; this article does not supply a fabricated server trace or claim a local speedup.
Then test retrieval, separately. Put an answer-bearing sentence near the beginning of one fixture and near the end of another. Query for each fact and inspect which chunk wins. This catches ingestion and retrieval mistakes together, but a low rank alone does not prove truncation. Prefix handling, chunk boundaries, and the embedding model can also cause it.
Ollama exposes prompt_eval_count, the number of input tokens processed. Keep it with batch size and timing while debugging. An aggregate batch count is not a per-document audit trail, and a positive value is not proof that every source paragraph reached the model. Individual probes make an ambiguous batch easier to inspect.
Repair the index without changing the wrong model
If an existing collection was built with truncation allowed, changing one request flag only protects future writes. It does not repair old vectors. Start with a sample of the longest stored chunks, check the original ingestion settings, and identify which records need re-embedding. If you cannot establish their provenance, rebuild into a new collection rather than guessing which rows are safe.
Store document ID, chunk boundaries, model tag or digest, prefix convention, dimension, and ingestion configuration alongside the indexing job. Preserve the exact source text separately. These fields let you explain why two vectors differ and avoid mixing embeddings created under incompatible configurations.
Build the replacement collection without overwriting the live one. Run the same retrieval questions against both. Include questions whose answers sit at document ends, exact identifiers that semantic search often misses, and ordinary user paraphrases. Promote the new collection only after it preserves the required matches. A larger accepted chunk is not automatically a better retrieval unit; it can dilute the answer among unrelated paragraphs.
Before paying for this repair, confirm you need the semantic branch. Our guide to starting with lexical retrieval explains when SQLite FTS5 is enough. If the correct chunk already ranks well but the generated answer drops its last paragraph, investigate the generator's prompt budget instead. That is the separate problem covered in our runtime diagnosis for local models.
Ollama recommends the same embedding model for indexing and querying. Keep that agreement explicit, including model-specific prefixes and any selected output dimensions. Replacing the generator cannot repair an index whose vectors never represented the answer-bearing text. Make rejected chunks visible first; then decide whether a different embedding model earns a rebuild.
Sources
- Ollama embed API: current input and response shape, default truncation, and error-on-overflow option.
- Ollama Nomic library: deployed v1.5 tag metadata and 2K context label.
- Nomic v1.5 model card: 8192 sequence length and required document/query task prefixes.
- Ollama embeddings guide: batching, normalized vectors, and consistent indexing/query models.
- Ollama issue 14186: a version-specific overflow report, not a universal statement about current behavior.