PyTorch / PyTorch Geometric dataloader
cvcdocdb.torch_dataloader streams a GraphStore
graph into PyTorch/PyTorch Geometric tensors without materializing the whole
graph in memory, for use with node embedding models (e.g. MetaPath2Vec)
and link-prediction pipelines.
Quick-start
from cvcdocdb import NetworkXGraph
from cvcdocdb.torch_dataloader import GraphDataset, to_hetero_edge_index_dict
graph = NetworkXGraph("bibliography.pkl")
dataset = GraphDataset(graph)
edge_index_dict = to_hetero_edge_index_dict(dataset)
Notes
GraphDataset/SubgraphDatasetstream nodes and edges in chunks rather than loading the entire graph into memory at once.label_filter/property_filteraccept the same MongoDB-style syntax ascvcdocdb.migration.migrate().split_edges_by_node_propertypartitions edges into N groups based on a scalar node property (e.g. train/test splits by anumthreshold, or more than two groups), and works for any property, not just numeric ones.edge_embeddingsderives an edge-level embedding from its endpoint node embeddings for use in link-prediction models.
See also the “PyTorch / PyTorch Geometric dataloader” tutorial notebook for a worked example on the bibliographic dataset (MetaPath2Vec training and link prediction).