When to use this
You needSearchType.TRIPLET_COMPLETION with recall() to return results. This search type matches queries against text representations of graph triplets (source β relationship β target).
Most users should start with improve(), which runs Cogneeβs default enrichment pass. Use this lower-level Memify pipeline directly when you specifically want to build or rebuild triplet embeddings.
Before you start:
- Complete Quickstart to understand basic operations
- Ensure you have LLM Providers configured
- Have an existing knowledge graph created with
remember()
Code in Action
What Just Happened
- Remember β builds a knowledge graph with entities and relationships from your text.
get_default_user()β retrieves the authenticated user. This pipeline requires aUserobject with write access to the dataset.create_triplet_embeddings(user, dataset)β iterates over all graph triplets, converts each to a text representation, and indexes them in the vector DB.- Recall with
TRIPLET_COMPLETIONβ queries the newTriplet_textcollection by semantic similarity.
What Changed in Your Graph
- The
Triplet_textcollection is populated in the vector DB. Each entry is a text representation of a graph triplet (source β relationship β target). recall(..., query_type=SearchType.TRIPLET_COMPLETION)queries now return results by matching your query against these triplet embeddings.
Additional Information
Parameters
Parameters
user(User, required) β authenticated user with write access. Obtain viaawait get_default_user().dataset(str, default:"main_dataset") β the dataset whose graph triplets to index.run_in_background(bool, default:False) β run asynchronously and return immediately.triplets_batch_size(int, default:100) β how many triplets to index per batch. Lower values use less memory; higher values are faster. The value does not affect coverage: triplets are paginated in a stable order, so a full pass visits each triplet exactly once at any batch size.
Under the hood
Under the hood
Two tasks run in sequence:
get_triplet_datapointsβ iterates over graph triplets and yieldsTripletobjects with embeddable text.index_data_pointsβ indexes each triplet in the vector DB under theTriplet_textcollection.
get_triplet_datapoints walks the whole graph with an offset loop, calling the graph adapterβs get_triplets_batch(offset, limit) with limit=triplets_batch_size and advancing the offset until a batch comes back short or empty. Every built-in adapter that supports batched reads sorts the triplets on (source node id, target node id, relationship name) before applying the offset, so consecutive calls slice one stable sequence and the loop cannot skip or repeat a triplet. Changing triplets_batch_size changes how the triplets are grouped into batches, not which triplets are visited.Troubleshooting
Troubleshooting
- Empty results from
TRIPLET_COMPLETIONβ ensure the graph has been built withremember()and thatcreate_triplet_embeddingsfinished without errors. - Some triplets missing from
Triplet_texton Neo4j or Ladybug β on releases before the pagination fix, those two adapters appliedSKIP/LIMITto an unordered result set, so a multi-batch run could skip triplets and leaveTriplet_textincomplete even though the pipeline reported success. Upgrade and re-runcreate_triplet_embeddingsto index the missed triplets. - Error: no graph data found β run
cognee.remember(..., dataset_name=...)before calling this pipeline. - LLM errors β verify that your LLM provider is configured. See LLM Providers.
- Permission errors β the user must have write access to the target dataset. See Permissions.
Improve
Understand the current improvement workflow
Self-Improvement Quickstart
Bridge session memory and enrich a dataset
Recall
Query the enriched graph with v1.0 retrieval