ocient_graph module brings a programming model (similar to Apache® Spark™ GraphX) to the Ocient® System directly from Python® using the pyocient driver. The module treats a graph as two relational tables, one for vertices (nodes) and one for edges (directed links). This module provides a composable API for graph transformations, neighborhood analytics, and iterative algorithms (e.g., Pregel, PageRank).
The API validates inputs, avoids destructive changes by materializing results into new tables, supports optional indexing for performance, and follows Ocient SQL conventions. The package installs separately from pyocient and exposes a Python-native interface that mirrors the Java® library. For details, see OCGraph Java Library.
Installation
Usepyocient for connectivity and ocient_graph for graph APIs. The graph library is a separate package that depends on pyocient. For a tutorial about installing and using pyocient, see Ocient Python Module: pyocient.
Install and Import
Install the ocient_graph module.
Shell
Python
Data Model Requirements
Database tables that use the OCGraph Python library must adhere to this structure. In addition to the listed requirements, tables can include other columns.Execution Mode
The OCGraph Python library supports two execution backends for graph algorithms: legacy and data flow. All algorithms prefer the data flow backend when the server supports it. You can override this default globally using theset_execution_config function.
The ExecutionMode enum defines the available modes.
Syntax
Python
Python
Python
Subgraph and Filtering
Use a subgraph or various filters to restrict a graph to relevant vertices and edges. These functions create filtered copies or masked intersections, preserving schema and optional indexes for performance.subgraph
Creates filtered vertex and edge tables using vertex and triplet predicates, retaining only edges with endpoints that remain after vertex filtering. The function creates the requested indexes and performs best-effort cleanup in the event of failure. SyntaxPython
Example
Create an active customer subgraph that includes only purchases exceeding $50 where the source and destination share a region.
Python
filter_vertices
Creates a filtered subgraph by selecting vertices that match a predicate while retaining only edges with endpoints that are in the filtered vertex set. SyntaxPython
Example
Filter US customers and retain edges with endpoints that remain in the filtered vertex set.
Python
filter_edges
Creates a filtered edges table by selecting edges that match a predicate. SyntaxPython
Example
This example demonstrates how to create a filtered edges table from an existing
purchases edge set by keeping only edges that meet a business rule (weight > 0.5 and ACTIVE). Then, the filterEdges method indexes the result on the srcid and destid columns for faster lookups.
Python
mask
Creates a masked subgraph by intersecting two graphs. Vertices intersect if the vertex identifier is present in both graphs. Edges intersect when thesrcid and destid values are present in both graphs.
The function creates a masked subgraph from rows that intersect with each other. The function copies rows that intersect from the graph defined by the arguments input_vertices_table and input_edges_table, including any attributes.
You can optionally create indexes on the result subgraph tables.
Syntax
Java
Example
Create a masked subgraph by intersecting two graphs. The example copies vertices and edges that are present in both graphs, along with the remaining endpoints.
Java
Transformations
Construct new vertex or edge tables by computing derived columns, reversing direction, or aggregating duplicates. These functions do not change the original inputs. Instead, the functions materialize new results.map_vertices
Creates a new vertices table with the identifierid and computed columns. Use the result_column_expressions argument to calculate additional columns. This function can also add indexes before inserting data.
Syntax
Python
Example
Create a new vertices table with two new columns,
name_upper and is_vip, and generate indexes for the id and name_upper columns.
Python
map_edges
Creates a new edges table withsrcid, destid, and any additional computed columns. Expressions should refer to input edge columns by their original names, and each computed expression should include an AS alias keyword.
Syntax
Python
Example
Create a new edges table with two columns,
discounted_amount and big_txn, and generate indexes for the srcid and destid columns.
Python
map_triplets
Creates a new edges table with computed columns that referencea (source vertex), b (edge), and c (destination vertex). The output automatically includes b.srcid and b.destid columns.
Syntax
Python
Example
Create a new triplet table from the vertices and edges tables with the
amount and same_country columns, and generate indexes for the src and destid columns.
Python
reverse_edges
Creates a new edges table with thesrcid and destid columns reversed, preserving other columns. Use this function to traverse a graph in the opposite direction.
Syntax
Python
Example
Transform edge direction by reversing the
srcid and destid columns. The example also creates indexes for these columns.
Python
group_edges
Groups duplicate rows of thesrcid and destid columns, producing one row for each unique pair of values in a new edges table. This function performs aggregations based on one or more SQL expressions.
Syntax
Python
Example
Create a new edge table that includes SQL aggregations for counting unique transactions
txn_count and total sums total_amount. Also, this function generates indexes for the src and destid columns.
Python
Triplets
Produce triplet representations that are made ofa (source vertex), b (edge), and c (destination vertex), either as a logical view or a materialized table for downstream queries.
create_triplets_view
Creates a view that combines the edge table with the source and destination vertex attributes. This view is useful for analyzing relationships without having to repeatedly join tables. The view includes these columns:- All original edge columns (including the
srcidanddestidcolumns). - All source-vertex columns except
id. Source-vertex column names have thesrc_prefix. - All destination-vertex columns except
id. Destination-vertex column names have thedest_prefix.
Python
Example
Create a triplets view to inspect edges with joined source and destination vertex attributes.
Python
create_triplets_table
Creates a materialized table that combines the edge table with the source and destination vertex attributes. This table is useful for analyzing relationships without having to repeatedly join tables. The created table includes these columns:- All original edge columns (including the
srcidanddestidcolumns). - All source-vertex columns except
id. Source-vertex column names have thesrc_prefix. - All destination-vertex columns except
id. Destination-vertex column names have thedest_prefix.
Python
Example
Create a new table for triplets. Generate indexes for the
src_id and dest_id columns.
Python
Degrees
Compute degree metrics for each vertex from the edges table. These functions produce small vertex tables suitable for joins and analytics.in_degrees
Computes how many edges point to each vertex in an edge table by counting how many times each uniquedestid value appears. The result table has two columns: id (the destination vertex) and in_degree (the count).
Syntax
Python
Example
Compute in-degrees per vertex and generate an index on the
id column.
Python
out_degrees
Computes how many edges originate from each vertex in an edge table by counting how many times each uniquesrcid value appears. The result table has two columns: id (the source vertex) and out_degree (the count).
Syntax
Python
Example
Compute the out-degrees count for each vertex and generate an index on the
id column.
Python
degrees
Computes the total degrees (in-degrees and out-degrees) for each vertex in an edge table by counting how many times each uniquesrcid and destid value appears. The result table has two columns: id (the destination or source vertex) and degree (the count).
Syntax
Python
Example
Compute total degrees for each vertex and generate an index on the
id column.
Python
Vertex Extraction and Joins
Build vertex sets from edges and combine vertex attributes across tables. These functions are useful for shaping vertex properties and consolidating features.from_edges
Builds a vertices table from an edges table by extracting the unique source and destination identifiers. This function can optionally compute additional columns using SQL expressions by referencing the unique identifier asids.id.
The created table always contains the id column with one additional column per expression.
Syntax
Python
Example
Create a vertices table from edge endpoints and add a
bucket column that assigns each vertex to one of 10 buckets. Generate an index for the id and bucket columns.
Python
join_vertices
Merges two vertices tables by retaining every row from a primary table (input_vertices_table) and selectively updating rows that also appear in the modification table (modification_vertices_table). The merged table includes all vertices from the primary table that do not appear in the modification table.
For vertices that appear in both tables, the function must include a list of expressions (resultAttributeExpressions) in the same column order for every non-identifier column in the merged result table. These SQL expressions can add computations to columns, or simply add aliases if no changes are needed. Each expression can reference columns from the primary table (using alias a) or from the modification table (using alias b).
Syntax
Python
Example
Merge vertex attributes and generate indexes for the
id and status columns. This example includes two SQL expressions to update the status and score columns based on the modification vertex table using the COALESCE SQL reference function.
Python
inner_join_vertices
Performs an inner join on two vertex tables using an equality comparisona.id = b.id. The result table automatically includes the id column from the first table.
The function must include a list of SQL expressions (result_attribute_expressions) in the same column order for every non-identifier column in the merged result table. These SQL expressions can add computations to columns, or simply add aliases if no changes are needed. Each expression can reference columns from the primary table using the alias a or from the modification table using the alias b.
Syntax
Python
Example
Create an inner join between two vertex tables and generate an index on the
id column.
Python
outer_join_vertices
Performs a left outer join between two vertices tables using an equality comparisona.id = b.id. The result table includes all rows from the left table. For left-table rows that have no match in the right table, any expression that reads columns from the right table with the alias b evaluates to NULL (while expressions that only read the table with the alias a remain non-NULL as usual).
The method must include a list of SQL expressions (result_attribute_expressions) in the same column order for every non-identifier column in the merged result table. These SQL expressions can add computations to columns, or simply add aliases if no changes are needed. Each expression can reference columns from the primary table using the alias a or from the modification table using the alias b.
Syntax
Python
Example
Perform a left outer join on two vertices tables and generate an index on the
id column.
Python
collect_neighbors
For each vertex in a table, this function collects information on neighbors (identifier and any attributes) as an array of tuples. For a specified direction (IN, OUT, or BOTH), the function aggregates tuples representing each neighboring vertex into an array.
The direction types are:
IN— Neighbors with edges pointing to the vertex (edges wheredestid = id).OUT— Neighbors that the vertex points to (edges wheresrcid = id).BOTH— Union ofINandOUTwith neighbors from incoming (destid = id) and outgoing (srcid = id) edges.
id (the vertex identifier) and neighbors (an array of tuples representing each neighbor).
If an error occurs after table creation, the function drops the result table.
Syntax
Python
Example
Collect incoming neighbors for each vertex and generate an index on the
id column. The direction argument set to IN collects neighbors pointing to id.
Python
collect_edges
For each vertex in a table, this function collects an array of adjacent edge rows based on the specified direction. The result table has two columns:id (the vertex identifier) and edges (an array of tuples, each tuple containing all columns from the edges table for a connected edge).
The direction types are:
IN— Edges pointing to the vertex (edges wheredestid = id).OUT— Edges originating from the vertex (edges wheresrcid = id).BOTH— Union ofINandOUTthat includes edges from incoming (destid = id) and outgoing (srcid = id) directions. This direction retains duplicates.
Python
Example
Collect outgoing edges per vertex. The example sets the
direction to OUT to collect edges from id.
Python
Algorithms
High-level graph algorithms that iterate over the graph structure to produce labels, components, or counts.label_propagation
Executes the Label Propagation Algorithm (LPA) to assign community labels to vertices. Each vertex starts with its own identifier as its label. For a number set by themaxIterations argument, each vertex updates its label to the most frequent label among its neighbors. The algorithm determines ties by choosing the smallest label. The algorithm uses temporary tables for intermediate results and drops these tables when the process completes or if it fails. Isolated vertices retain their initial label. The final table stores id and label columns and can include indexes.
Syntax
Python
Example
Run label propagation for 10 iterations and assign labels to vertices. Generate an index on the
id column.
Python
connected_components
Identifies the connected components of an undirected graph. This algorithm configures a Pregel computation in which each vertex initially sets its component label equal to its own identifierid.
In each iteration, vertices send their component label to neighbors. Each vertex updates based on the aggregated minimum value of its current component label and any received values. The process repeats until no more updates occur.
The result table maps each vertex id to its final component label.
Syntax
Python
Example
Compute connected components and generate an index on the
id column.
Python
strongly_connected_components
Computes strongly connected components (SCC) in a directed graph. This function runs a recursive algorithm that partitions vertices into subsets where every vertex is reachable from other vertices in the same subset. This function uses recursive partitioning. The algorithm selects a pivot (typically the minimum identifierid), computes its predecessor set (vertices that can reach the pivot), and its descendant sets (vertices reachable from the pivot). Then, the function identifies the SCC as their intersection, removes that SCC from the graph, and recurses on the remainder until all vertices have been assigned to an SCC. The output contains columns for the id and component identifiers (the minimum id in the SCC).
The function creates temporary tables in the result schema to store intermediate results. This function drops these tables when the computation completes or fails. The final result table contains two columns: id (vertex identifier) and component (the minimum vertex identifier in its SCC subset).
Syntax
Python
Example
Compute the SCC and generate an index on the
id column.
Python
TriangleCount
TriangleCount identifies all 3-cycles (triangles) in the graph and counts how many distinct triangles each vertex participates in. The algorithm first builds a canonical, undirected edge set by ensuringsrcid < destid and removing duplicates to prevent double-counting. If your input edges are already canonicalized and deduplicated, use TriangleCount.run_pre_canonicalized to skip preprocessing for faster performance.
The function then counts triangles (a, b, c) where a < b < c by intersecting neighbor lists and aggregates per-vertex participation to produce a result table with the id and triangle_count columns.
run syntax
Python
run_pre_canonicalized syntax
Python
Examples
Count Triangles Using
run
Canonicalize the raw edges internally, count unique triangles, and write per-vertex triangle counts with an index on the id column.
Python
run_pre_canonicalized
Use a pre-canonicalized, deduplicated edge table to count triangles and write per-vertex triangle counts with an index on the id column.
Python
pregel
Provides a generic vertex‑centered iteration framework for custom graph algorithms, similar to the Pregel model. Each iteration updates vertex states by sending messages along edges and then aggregating these messages to compute new states. The algorithm continues iterating until it reaches convergence (no state changes or no messages produced) or a specified iteration cap. The algorithm uses multiple specified SQL expressions. SyntaxPython
Example
Run a simple Pregel computation summing incoming edge amounts into the vertex state for 10 iterations at most, and generate an index on the
id column.
Python
Paths & Ranking
These functions include the shortest-path and PageRank algorithms.shortest_paths
Computes the shortest distance from every vertex to each set of landmark vertices using an iterative relaxation algorithm. The algorithm resembles Bellman–Ford but simultaneously handles multiple destinations. Each landmark starts at distance0 and all others at positive infinity. On each iteration, the algorithm examines every edge and checks whether traveling through the connected neighbor would yield a shorter route to a landmark. If a shorter route exists, the algorithm updates the distance of the source vertex. The process stops when no distances improve or the algorithm reaches the maximum number of iterations.
After the process finishes, the algorithm writes a result table with the srcid, destid, and distance columns.
Syntax
Python
Example
Compute distances from landmarks and generate indexes on the
src and dest columns.
Python
static_page_rank
Computes PageRank scores for each vertex over a fixed number of iterations. The algorithm follows the standard PageRank formula with a damping factor (damping_factor) and uses common table expressions to calculate contributions from incoming edges and redistribute ranks from dangling nodes.
The algorithm supports two variants:
- Standard PageRank — All vertices start with rank
1.0/N, whereNis the number of vertices. Specify this variant ifpersonalizationSrcIdisnull. - Personalized PageRank — The specified vertex starts with a rank of
1.0, while others start with a rank of0.0. Specify this variant ifpersonalizationSrcIdis a vertex identifier.
Python
Example
Run fixed-iteration PageRank and generate an index on the
id column. This example uses a damping factor of 0.85 to ensure the ranking concentrates on highly linked regions.
Python
dynamic_page_rank
Computes PageRank scores until convergence based on a specified threshold value (tolerance). Unlike the static_page_rank function, this algorithm runs iterations until the sum of absolute differences between ranks in successive iterations is less than or equal to the tolerance value. The algorithm handles personalization similarly to static_page_rank. At each iteration, the function uses the PageRank formula, collects rank values, and redistributes them.
The algorithm supports two variants:
- Standard PageRank — All vertices start with rank
1.0/N, whereNis the number of vertices. Specify this variant ifpersonalizationSrcIdisnull. - Personalized PageRank — The specified vertex starts with a rank of
1.0, while others start with a rank of0.0. Specify this variant ifpersonalizationSrcIdis a vertex identifier.
tolerance threshold, the function writes a vertices table containing all the original vertex columns with a new PageRank scoring column.
Syntax
Python
Example
Run dynamic PageRank to convergence and generate an index on the
id column. This example uses a low tolerance value of 1.0e-6, which generates high-precision rankings but requires more computing resources.
Python
You can specify the deprecated
reset_prob keyword argument instead of the damping_factor argument. The value means the same thing (the damping factor, and not the teleport probability). Specifying the reset_prob argument raises the DeprecationWarning warning. You cannot specify both damping_factor and reset_prob arguments in the same function call, but you must specify one of them.stable_marriage
Computes a stable matching between two groups using the Gale–Shapley algorithm. Each suitor proposes to candidates in order of preference, and each candidate tentatively holds the best proposal received. The algorithm iterates until no suitor can improve or until it reachesmax_iterations rounds, producing a suitor-optimal stable matching.
The suitor preferences table must have the suitor_id, candidate_id, and rank columns. The candidate preferences table must have the candidate_id, suitor_id, and rank columns. All rank values must be non-NULL. The result table contains the suitor_id and candidate_id columns.
Syntax
Python
Example
Compute a stable matching between suitors and candidates with a maximum of 20 rounds.
Python
jaccard_similarity
Computes the Jaccard similarity for every pair of vertices that share at least one neighbor. The Jaccard similarity between two vertices is the size of the intersection of their neighbor sets divided by the size of their union. This metric is useful for link prediction and duplicate detection. The result table contains thesrcid, destid, and similarity columns.
Syntax
Python
Example
Compute pairwise Jaccard similarity scores.
Python
cosine_similarity
Computes the cosine similarity for every pair of vertices that share at least one neighbor. The cosine similarity represents the cosine of the angle between two neighbor-set vectors. This metric is useful for recommendation systems and measuring structural equivalence. The result table contains thesrcid, destid, and similarity columns.
Syntax
Python
Example
Compute pairwise cosine similarity scores.
Python
k_core_decomposition
Computes the k-core decomposition of a graph by iteratively removing vertices with degree less thank until only the maximal subgraph with the minimum degree k remains. The algorithm computes the coreness value for each vertex.
The result table contains the id and core columns.
Syntax
Python
Example
Compute k-core decomposition and index on the
id column.
Python
eigenvector_centrality
Computes eigenvector centrality scores for each vertex by iteratively updating the score of each vertex to be the sum of the scores of its neighbors, followed by normalization. Vertices connected to other high-scoring vertices receive higher centrality. The result table contains theid and centrality columns.
Syntax
Python
Example
Compute eigenvector centrality with 20 iterations.
Python
cycle_detection
Detects vertices that participate in cycles in a directed graph using iterative degree-based peeling. The algorithm repeatedly removes vertices with an in-degree or out-degree of zero until no such vertices remain. The remaining vertices are those involved in at least one cycle. The result table includes theid column for each vertex in a cycle.
Syntax
Python
Example
Detect vertices involved in cycles.
Python
max_bipartite_matching
Computes maximum matching in a bipartite graph using an iterative augmenting-path approach. The algorithm greedily matches unmatched vertices and then refines the matching until it finds no further augmenting path or until it reaches the number of rounds as specified by themax_iterations value. The result is a set of edges where no two edges share a vertex.
The result table contains the srcid and destid columns.
Syntax
Python
Example
Compute a maximum bipartite matching with up to 10 augmentation rounds.
Python
louvain_modularity_optimization
Detects communities using the Louvain method, a greedy modularity optimization algorithm. The algorithm iteratively moves vertices between communities to maximize modularity, then aggregates communities and repeats until modularity stops improving or until it reaches the number of rounds as specified by themax_iterations value.
The result table contains the id and community columns.
Syntax
Python
Example
Detect communities using the Louvain method with a maximum of 10 iterations on an unweighted graph.
Python
aggregate_messages
A general-purpose message-passing primitive that sends messages along edges and aggregates them at destination vertices. This function is useful for building custom graph computations, feature engineering, and implementing algorithms not available as built-in functions. For each edge, the function evaluates a message expression referencing source vertex attributes (a.*), edge attributes (b.*), and destination vertex attributes (c.*). The function then aggregates the messages for each destination vertex using a SQL aggregation function.
The result table contains the vertex id and the aggregated message column, named using the AS alias in the aggregate_expr argument.
Syntax
Python
Example
Compute the sum of incoming edge weights for each vertex.
Python
a_star_shortest_path
Computes the shortest path between a source and target vertex using the A* search algorithm. A* extends Dijkstra’s algorithm with a specified heuristic expression that estimates the remaining cost to the target, guiding the search toward the goal, reducing the number of vertices explored, and stopping early if it reaches the number of iterations as specified by themax_iterations value.
The heuristic must be admissible, meaning it never overestimates the true cost, for the algorithm to guarantee an optimal path. The result table contains the srcid, destid, cost, and path_index columns, which represent the ordered nodes on the shortest path.
Syntax
Python
Example
Find the shortest path from vertex
1 to vertex 42 using edge weights and a coordinate-based heuristic.
Python
yens_k_shortest_paths
Computes thek shortest loopless paths between a source and target vertex using Yen’s algorithm. The algorithm iteratively finds the next shortest path by deviating from previously discovered paths at each spur node, using up to the number of Dijkstra iterations as specified by the max_iterations value for each spur-node search.
The result table contains the path_id, srcid, destid, cost, and path_index columns, which represent the ordered nodes on each path.
Syntax
Python
Example
Find the three shortest paths from vertex
1 to vertex 42 using edge weights.
Python
leiden_modularity_optimization
Detects communities using the Leiden algorithm, a refinement of the Louvain method that guarantees the formation of well-connected communities. The algorithm optimizes modularity through iterative local moves and a refinement phase that prevents poorly connected communities, repeating until modularity stops improving or until it reaches the number of rounds as specified by themax_iterations value.
The result table contains the id and community columns.
Syntax
Python
Example
Detect communities using the Leiden algorithm with a resolution of
1.0 and a maximum of 10 iterations.
Python
time_respecting_shortest_path
Computes the shortest path between a source and target vertex while respecting the temporal ordering of edges. Each edge has a timestamp, and the algorithm traverses only edges whose timestamps are nondecreasing along the path, modeling real-world scenarios where events must occur in chronological order. The timestamp column must be ofTIMESTAMP type. The result table contains the srcid, destid, cost, and path_index columns.
Syntax
Python
Example
Find the time-respecting shortest path from vertex
1 to vertex 42 departing no earlier than midnight on January 1, 2024.
Python
temporal_page_rank
Computes PageRank scores using exponential time decay on edges. Recent edges contribute more to the ranking than older edges, making this algorithm suitable for graphs where recency matters (e.g., communication networks, transaction logs). The time column must be ofTIMESTAMP type. The algorithm supports both standard and personalized variants, similar to static_page_rank.
Syntax
Python
Example
Run temporal PageRank with a decay factor of
0.5 using a reference time of December 31, 2024.
Python
random_walk
Generates configurable random walks from each vertex using Node2Vec-style biased sampling. The return parameter (p_return argument) and in-out parameter (q_inout argument) control whether the walk favors revisiting the previous node (breadth-first) or exploring further (depth-first).
The result table contains the walk_id, step, and node_id columns.
Syntax
Python
Example
Generate five random walks of length 10 from each vertex using Node2Vec-style parameters.
Python
fast_random_projection
Generates node embeddings using the Fast Random Projection (FastRP) algorithm. FastRP creates low-dimensional vector representations of vertices by iteratively averaging neighbor embeddings with sparse random projections. These embeddings are useful as features for downstream machine learning tasks. The result table contains theid and embedding columns. The embedding column is a vector with length equal to the embedding_dim value.
Syntax
Python
Example
Generate 64-dimensional normalized embeddings with three iterations.
Python

