DRM Tools Documentation
Document Representation Model (DRM) is a Python library for graph-based document representation with Neo4j and an in-memory NetworkX backend.
Contents:
Features
Graph-based document representation: Model documents as nodes and relations, including hierarchical structures such as Document → Section → Page.
Two backends: Full Neo4j integration via
Neo4jGraphor an in-memoryNetworkXGraph(NetworkX) for testing and tutorials.Two entity levels: Root entities (
Node) and child entities (WeakNode) with composite primary keys and cascade delete propagation.Dependency auto-insertion: String properties (for example names) are automatically materialised as
Valornodes when the graph backend supports them.FK validation: Foreign key constraints on relations prevent dangling references.
ON DELETE strategies: CASCADE, RESTRICT, SET NULL.
Query & filtering: Secondary index on scalar properties, multi-filter search with intersection/union, and debug snapshots.
Vector search (NetworkX only): HNSW-based ANN indexing on node properties with
cosine,l2, andipdistance spaces.Propagation properties: Every node and edge can carry
_propagate,is_weak,parent_relation, and_dependenciesflags that enable cascade delete and hierarchical traversal.Transactional group creation:
create_group()creates a strong node together with its WeakNodes and WeakRelations in a single isolated transaction. On failure the entire group is rolled back.Lazy background initialization:
init_propagation()scans the backend graph, detects WeakNodes from edge structure, and initializes propagation properties. Supports background mode and progress callbacks.
Primary Key
Every node must have a primary key (pk) or a neo4j_id (or both).
The pk is either an int or a dict:
# Integer PK — converted to {"id": value}
node = Node(pk=42, main_label="Document")
# Dict PK — used as-is
node = Node(pk={"doc_id": "DOC-001", "version": 2}, main_label="Document")
# Auto-assigned PK — pass None explicitly; the backend generates an ID
node = Node(pk=None, main_label="AutoIdNode")
graph.insertNode(node) # node._primary_key becomes {"id": <generated_id>}
# Reconstructed from DB — neo4j_id is used as PK
node = Node(neo4j_id=123, main_label="Document") # _primary_key = {"id": 123}
Rules:
Node(pk={"id": 1}, main_label="X")→_primary_key = {"id": 1}Node(pk=42, main_label="X")→_primary_key = {"id": 42}Node(pk=None, main_label="X")→_primary_key = None(backend assigns after insert)Node(pk=None, neo4j_id=123, main_label="X")→_primary_key = {"id": 123}Node(main_label="X")→ ValueError — pk must be provided
A node with pk=None is valid — the backend (Neo4j or NetworkX) assigns
an auto-generated ID as the primary key after insertion. If no backend is
used, _primary_key remains None.
Query & Filtering
NetworkXGraph maintains a secondary index on scalar properties, enabling
fast lookups without scanning all nodes:
from cvcdocdb import NetworkXGraph, Node
graph = NetworkXGraph()
alice = Node(pk={"name": "Alice"}, main_label="Author")
alice["affiliation"] = "MIT"
bob = Node(pk={"name": "Bob"}, main_label="Author")
bob["affiliation"] = "Stanford"
graph.insertNode(alice)
graph.insertNode(bob)
# Find all authors from MIT
mit_authors = graph.find_nodes_by_property("affiliation", "MIT")
print(mit_authors) # [node_id_of_alice]
# Multi-filter search (intersection or union)
results = graph.find_nodes({"affiliation": "MIT"}, match="all")
# Debug snapshot
graph.print_debug()
Vector Indexing
The package exposes a common vector-index API on graph stores:
enable_vector_index(property_name, dimensions, space="cosine", ...)query_vector_index(property_name, vector, top_k=10)
Current backend support:
NetworkXGraph: supported (useshnswlib)Neo4jGraph: API is present but currently raisesNotImplementedError
Example (NetworkXGraph):
from cvcdocdb import NetworkXGraph, Node
graph = NetworkXGraph()
graph.enable_vector_index("embedding", dimensions=3, space="cosine")
graph.insertNode(Node(pk={"id": 1}, main_label="Doc", embedding=[1.0, 0.0, 0.0]), replace=True)
graph.insertNode(Node(pk={"id": 2}, main_label="Doc", embedding=[0.0, 1.0, 0.0]), replace=True)
results = graph.query_vector_index("embedding", [1.0, 0.0, 0.0], top_k=2)
print(results) # [(node_id, distance), ...]
graph.close()
Propagation Properties
Every node and edge in the graph can carry propagation-related properties. These properties come in two families: structural (describing the hierarchy) and operational (tracking initialization state).
Structural properties — describe the parent-child relationship:
_propagate(bool): marks edges that trigger cascade delete when the parent node is removed. Set automatically on WeakNode parent-child edges.is_weak(bool): marks child nodes (WeakNodes) that are linked to their parent by a_propagateedge. Use this to find all nodes that will be cascade-deleted when their parent is removed.
Operational properties — track initialization state:
parent_relation(str): stores the edge type linking a parent to its WeakNode child (e.g."HAS_SECTION","HAS_PAGE")._dependencies(dict): tracks auto-insertedValornodes for string properties._weak_init_done(bool): tracks whether a strong node (parent) has been processed byinit_propagation(). Set automatically bycreate_group()because the edges already carry_propagate=True.
Example:
from cvcdocdb import Neo4jGraph, Node, WeakNode
graph = Neo4jGraph(
url="bolt://localhost:7687",
user="neo4j",
password="secret",
database="mydb",
)
doc = Node(pk={"doc_id": "DOC-1"}, main_label="Document")
graph.insertNode(doc)
section = WeakNode(
parent=doc,
parent_relation="HAS_SECTION",
pk={"section": "intro"},
main_label="Section",
)
graph.insertNode(section, insert_parent=True)
# The edge doc → section carries _propagate=True
# The section node carries is_weak=True
# After init_propagation():
# section.is_weak = True (structural: it's a WeakNode)
# section._propagate = True (structural: cascade-delete enabled)
# section.parent_relation = "HAS_SECTION" (operational: edge type)
# doc._weak_init_done = True (operational: parent initialized)
Key distinction — is_weak vs _weak_init_done:
``is_weak`` lives on the child node and describes its role in the hierarchy. It is set by
init_propagation()when it detects a node connected via a_propagateedge.``_weak_init_done`` lives on the parent node and tracks whether that parent’s WeakNodes have been processed. It is set automatically by
create_group()because the group is created with edges that already carry_propagate=True.
# After create_group(doc, [section, page]):
doc._weak_init_done = True # parent: "my children are initialized"
section.is_weak = True # child: "I am a WeakNode"
page.is_weak = True # child: "I am a WeakNode"
# After init_propagation() on a pre-existing graph:
section.is_weak = True # child: "I was detected as a WeakNode"
doc._weak_init_done = True # parent: "I have been processed"
Transactional Group Creation
create_group() inserts a strong node together with its WeakNodes and
WeakRelations in a single atomic transaction. On failure the entire group
is rolled back so the graph is never left in a partially-created state.
Example:
from cvcdocdb import Neo4jGraph, Node, WeakNode
graph = Neo4jGraph(
url="bolt://localhost:7687",
user="neo4j",
password="secret",
database="mydb",
)
doc = Node(pk={"doc_id": "DOC-1"}, main_label="Document", title="My Document")
section = WeakNode(
parent=doc,
parent_relation="HAS_SECTION",
pk={"section": "intro"},
main_label="Section",
title="Introduction",
)
page = WeakNode(
parent=section,
parent_relation="HAS_PAGE",
pk={"page": 1},
main_label="Page",
content="Welcome to the document.",
)
# Atomic creation — all-or-nothing
doc_id = graph.create_group(
strong_node=doc,
weak_nodes=[section, page],
)
print(f"Created group with strong node id: {doc_id}")
# _weak_init_done is automatically set on the strong node
# because the edges already carry _propagate=True
Lazy Background Initialization
init_propagation() scans the backend graph and initializes propagation
properties on nodes and edges that are missing them. It detects WeakNodes
by looking for edges with _propagate=True and marks the child nodes
accordingly.
Example:
from cvcdocdb import Neo4jGraph
graph = Neo4jGraph(
url="bolt://localhost:7687",
user="neo4j",
password="secret",
database="mydb",
)
# Synchronous — blocks until complete
result = graph.init_propagation()
print(f"Initialized: {result}") # True on first call, False if already done
# Background mode — returns immediately, runs in daemon thread
graph.init_propagation(background=True)
# With progress callback
def progress(processed, total):
print(f" {processed}/{total} nodes processed")
graph.init_propagation(progress_callback=progress)
# Idempotent — second call returns False immediately
graph.init_propagation() # Returns False
Installation
Install the package in development mode:
pip install -e .
Quick Start
from cvcdocdb import Neo4jGraph, Node, WeakNode
# Connect to Neo4j
graph = Neo4jGraph(
url="bolt://localhost:7687",
user="neo4j",
password="secret",
database="mydb",
)
# Create a document hierarchy
doc = Node(pk={"doc": "DOC-001"}, main_label="Document")
graph.insertNode(doc, replace=True)
section = WeakNode(parent=doc, pk={"section": 1}, main_label="Section")
graph.insertNode(section, insert_parent=True)
page = WeakNode(parent=section, pk={"page": 1}, main_label="Page")
graph.insertNode(page, insert_parent=True)
graph.close()
For the in-memory backend:
from cvcdocdb import NetworkXGraph, Node
graph = NetworkXGraph()
doc = Node(pk={"doc": "DOC-001"}, main_label="Document")
graph.insertNode(doc)
print(graph.get_node_ids()) # [1]
graph.close()
Configuration
DRM uses environment variables for Neo4j connections. Multiple targets are
supported via the NEO4J_TARGET selector:
export NEO4J_DEV_URL=bolt://dev-host:7687
export NEO4J_DEV_USER=neo4j
export NEO4J_DEV_PASSWORD=your_dev_password
export NEO4J_DEV_DATABASE=neo4j
For a custom target:
export NEO4J_TARGET=LOCAL
export NEO4J_LOCAL_URL=bolt://localhost:7687
export NEO4J_LOCAL_USER=neo4j
export NEO4J_LOCAL_PASSWORD=your_password
export NEO4J_LOCAL_DATABASE=neo4j
The real Neo4j tests default to DEV when NEO4J_TARGET is not set,
and fall back to plain NEO4J_* variables when prefixed variables are absent.
Running Tests
Run the test suite with pytest:
python -m pytest test/ -v
Tests use NetworkXGraph for all unit tests and Neo4jGraph for
integration tests against a real Neo4j instance (controlled by NEO4J_TARGET).
Building Documentation
Generate the HTML documentation with Sphinx:
cd docs
sphinx-build -b html . _build/html
Or use the virtual environment:
.venv/bin/sphinx-build -b html docs/ docs/_build/html/
Acknowledgements
This work has been partially supported by the Spanish project PID2024-157778OB-I00, Ministerio de Ciencia e Innovación, the Departament de Cultura of the Generalitat de Catalunya, and the CERCA Program / Generalitat de Catalunya. Adrià Molina is funded with the PRE2022-101575 grant provided by MCIN / AEI / 10.13039 / 501100011033 and by the European Social Fund (FSE+).