> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ocient.com/llms.txt
> Use this file to discover all available pages before exploring further.

# OCGraph Python Library

> OCGraph Python provides a SQL-native interface for graph analytics on Ocient using a Python workflow, with function references, parameter tables, and examples.

export const Spark = "Apache® Spark™";

export const Python = "Python®";

export const Ocient = "Ocient®";

export const Java = "Java®";

export const IEEE = "IEEE®";

The `ocient_graph` module brings a programming model (similar to {Spark} GraphX) to the {Ocient} System directly from {Python} using the `pyocient` driver. The module treats a graph as two relational tables, one for vertices (nodes) and one for edges (directed links). This module provides a composable API for graph transformations, neighborhood analytics, and iterative algorithms (e.g., Pregel, PageRank).

The API validates inputs, avoids destructive changes by materializing results into new tables, supports optional indexing for performance, and follows Ocient SQL conventions. The package installs separately from `pyocient` and exposes a Python-native interface that mirrors the {Java} library. For details, see [OCGraph Java Library](/ocgraph-java-library).

## Installation

Use `pyocient` for connectivity and `ocient_graph` for graph APIs. The graph library is a separate package that depends on `pyocient`. For a tutorial about installing and using `pyocient`, see [Ocient Python Module: pyocient](/ocient-python-module-pyocient).

**Install and Import**

Install the `ocient_graph` module.

```shell Shell theme={null}
pip install ocient_graph
```

Import the module.

```python Python theme={null}
from pyocient import connect
from ocient_graph import (
    subgraph,
    collect_neighbors,
    EdgeDirection,
)
```

### Data Model Requirements

Database tables that use the OCGraph Python library must adhere to this structure. In addition to the listed requirements, tables can include other columns.

| **Table** | **Description** | **Requirements** |
| - | - | - |
| Vertices table | A table with one row per vertex (node). This table typically represents the anchor for graph algorithms and transforms. Many methods join edges to vertices by the `id` column. | The table must contain the<br />`id BIGINT NOT NULL` column definition as the unique vertex identifier. |
| Edges table | A table with one row per directed edge (relationship). Each row is a directed edge from a source vertex to a destination vertex. | The table must contain these column definitions:<br />`srcid BIGINT NOT NULL`,<br />`destid BIGINT NOT NULL` |

### Execution Mode

The OCGraph Python library supports two execution backends for graph algorithms: legacy and data flow. All algorithms prefer the data flow backend when the server supports it. You can override this default globally using the `set_execution_config` function.

The `ExecutionMode` enum defines the available modes.

| **Mode** | **Description** |
| - | - |
| `ExecutionMode.ALWAYS_LEGACY` | Forces all algorithms to use the legacy SQL execution backend. |
| `ExecutionMode.ALWAYS_DATAFLOW` | Forces all algorithms to use the data flow execution backend. If the server does not support data flows, the function raises a `DatabaseError` error. |
| `ExecutionMode.AUTO_FASTEST` | Selects the backend for each algorithm automatically, based on whether the server supports data flows. This mode is the default. |

**Syntax**

```python Python theme={null}
from ocient_graph import set_execution_config, ExecutionMode

set_execution_config(mode)
```

**Examples**

**Set the Data Flow Mode**

Force all algorithms to run in data flow mode.

```python Python theme={null}
set_execution_config(ExecutionMode.ALWAYS_DATAFLOW)
```

**Reset to the Default Mode**

Reset to the default automatic mode.

```python Python theme={null}
set_execution_config(ExecutionMode.AUTO_FASTEST)
```

## Subgraph and Filtering

Use a subgraph or various filters to restrict a graph to relevant vertices and edges. These functions create filtered copies or masked intersections, preserving schema and optional indexes for performance.

### subgraph

Creates filtered vertex and edge tables using vertex and triplet predicates, retaining only edges with endpoints that remain after vertex filtering. The function creates the requested indexes and performs best-effort cleanup in the event of failure.

**Syntax**

```python Python theme={null}
subgraph(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_vertices_table,
    result_edges_table,
    vertex_filter,
    edge_filter,
    [ result_vertices_indexes [ , ... ] ],
    [ result_edges_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the filtered vertices table to create. |
| `result_edges_table` | str | Name of the filtered edges table to create. |
| `vertex_filter` | str | SQL predicate to filter the vertices (without the `WHERE` keyword). Example: `status = 'ACTIVE' AND score > 0` |
| `edge_filter` | str | SQL predicate that the system evaluates in a triplet context using the aliases `a` (source vertex), `b` (edge), and `c` (destination vertex). This predicate does not require a `WHERE` keyword. <br />Example: `b.amount > 50 AND a.region = c.region` |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |
| `result_edges_indexes` | List\[str] | Optional. Columns to index in the result edges table (e.g., `srcid` and `destid` columns). Specify an empty list for none. |

**Example**
Create an active customer subgraph that includes only purchases exceeding \$50 where the source and destination share a region.

```python Python theme={null}
subgraph(
    connection,
    "sales",
    "customers",
    "purchases",
    "sales",
    "customers_active",
    "purchases_active",
    "status = 'ACTIVE' AND score > 0",
    "b.amount > 50 AND a.region = c.region",
    ["id","region"],
    ["srcid","destid"],
)
```

### filter\_vertices

Creates a filtered subgraph by selecting vertices that match a predicate while retaining only edges with endpoints that are in the filtered vertex set.

**Syntax**

```python Python theme={null}
filter_vertices(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_vertices_table,
    result_edges_table,
    vertex_filter,
    [ result_vertices_indexes [ , ... ] ],
    [ result_edges_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the result vertices table. The name must not conflict with the names of input tables. |
| `result_edges_table` | str | Name of the edges table to create. The name must not conflict with the names of input tables. |
| `vertex_filter` | str | SQL predicate to filter the vertices (without the `WHERE` keyword). Example: `status = 'ACTIVE' AND score > 0` |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |
| `result_edges_indexes` | List\[str] | Optional. Columns to index in the result edges table (e.g., `srcid` and `destid` columns). Specify an empty list for none. |

**Example**
Filter US customers and retain edges with endpoints that remain in the filtered vertex set.

```python Python theme={null}
filter_vertices(
    connection,
    "sales", "customers", "purchases",
    "sales", "customers_us", "purchases_us",
    "country = 'US'",
    ["id"],
    ["srcid","destid"],
)
```

### filter\_edges

Creates a filtered edges table by selecting edges that match a predicate.

**Syntax**

```python Python theme={null}
filter_edges(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_edges_table,
    edge_filter,
    [ result_edges_indexes [, ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty achema containing the input edges table. |
| `input_edges_table` | str | Input edges table (must have `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_edges_table` | str | Result edges table name. The name must not conflict with the names of input tables. |
| `edge_filter` | str | SQL predicate on edges (without the `WHERE` keyword). Example: `weight > 0.5 AND type = 'ACTIVE'` |
| `result_edges_indexes` | List\[str] | Optional. Columns to index in the result edges table (e.g., `srcid` and `destid` columns). Specify an empty list for none. |

**Example**
This example demonstrates how to create a filtered edges table from an existing `purchases` edge set by keeping only edges that meet a business rule (`weight > 0.5` and `ACTIVE`). Then, the `filterEdges` method indexes the result on the `srcid` and `destid` columns for faster lookups.

```python Python theme={null}
filter_edges(
    connection,
    "sales", "purchases",
    "sales", "purchases_filtered",
    "weight > 0.5 AND type = 'ACTIVE'",
    ["srcid", "destid"],
)
```

### mask

Creates a masked subgraph by intersecting two graphs. Vertices intersect if the vertex identifier is present in both graphs. Edges intersect when the `srcid` and `destid` values are present in both graphs.

The function creates a masked subgraph from rows that intersect with each other. The function copies rows that intersect from the graph defined by the arguments `input_vertices_table` and `input_edges_table`, including any attributes.

You can optionally create indexes on the result subgraph tables.

**Syntax**

```java Java theme={null}
mask(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    other_schema,
    other_vertices_table,
    other_edges_table,
    result_schema,
    result_vertices_table,
    result_edges_table,
    [ result_vertices_indexes [ , ... ] ],
    [ result_edges_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | The primary vertices table to intersect (must have an `id` column).<br /><br />The masked subgraph created by this function copies rows that intersect from this table. |
| `input_edges_table` | str | The primary edges table to intersect (must have `srcid` and `destid` columns).<br /><br />The masked subgraph created by this function copies rows that intersect from this table. |
| `other_schema` | str | Schema containing the second graph. |
| `other_vertices_table` | str | The second vertices table to intersect. |
| `other_edges_table` | str | The second edges table to intersect. |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create. |
| `result_edges_table` | str | Name of the edges table to create. |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |
| `result_edges_indexes` | List\[str] | Optional. Columns to index in the result edges table (e.g., `srcid` and `destid` columns). Specify an empty list for none. |

**Example**

Create a masked subgraph by intersecting two graphs. The example copies vertices and edges that are present in both graphs, along with the remaining endpoints.

```java Java theme={null}
mask(
    connection,
    "sales", "customers", "purchases",
    "ref",   "customers_ref", "purchases_ref",
    "sales", "customers_masked", "purchases_masked",
    ["id"],
    ["srcid", "destid"],
)
```

## Transformations

Construct new vertex or edge tables by computing derived columns, reversing direction, or aggregating duplicates. These functions do not change the original inputs. Instead, the functions materialize new results.

### map\_vertices

Creates a new vertices table with the identifier `id` and computed columns. Use the `result_column_expressions` argument to calculate additional columns. This function can also add indexes before inserting data.

**Syntax**

```python Python theme={null}
map_vertices(
    connection,
    input_schema,
    input_vertices_table,
    result_schema,
    result_vertices_table,
    result_column_expressions,
    [ result_vertices_indexes [, ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input vertices table. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `result_schema` | str | A schema to create the result vertices table. |
| `result_vertices_table` | str | Name of the result vertices table. The name must not conflict with the names of input tables. |
| `result_column_expressions` | List\[str] | One or more SQL expressions defining result columns beyond `id`. Use the `AS alias_name` keyword for stable names. |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Create a new vertices table with two new columns, `name_upper` and `is_vip`, and generate indexes for the `id` and `name_upper` columns.

```python Python theme={null}
map_vertices(
    connection,
    "sales", "customers",
    "sales", "customers_enriched",
    [
      "UPPER(name) AS name_upper",
      "CASE WHEN score > 1000 THEN true ELSE false END AS is_vip",
    ],
    ["id","name_upper"],
)
```

### map\_edges

Creates a new edges table with `srcid`, `destid`, and any additional computed columns. Expressions should refer to input edge columns by their original names, and each computed expression should include an `AS alias` keyword.

**Syntax**

```python Python theme={null}
map_edges(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_edges_table,
    [ result_column_expressions [, ... ] ],
    [ result_edges_indexes [, ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input edges table. |
| `input_edges_table` | str | Input edges table (must have `srcid` and `destid` columns). |
| `result_schema` | str | A schema to create the result edges table. |
| `result_edges_table` | str | Name of the result edges table. The name must not conflict with the names of input tables. |
| `result_column_expressions` | List\[str] | Optional. SQL expressions for additional edge columns (besides the `srcid` and `destid` columns). Use the `AS alias_name` keyword for stable names. |
| `result_edges_indexes` | List\[str] | Optional. Columns to index in the result edges table (e.g., `srcid` and `destid` columns). Specify an empty list for none. |

**Example**

Create a new edges table with two columns, `discounted_amount` and `big_txn`, and generate indexes for the `srcid` and `destid` columns.

```python Python theme={null}
map_edges(
    connection,
    "sales", "purchases",
    "sales", "purchases_enriched",
    [
      "amount * 0.9 AS discounted_amount",
      "CASE WHEN amount > 100 THEN 1 ELSE 0 END AS big_txn",
    ],
    ["srcid","destid"],
)
```

### map\_triplets

Creates a new edges table with computed columns that reference `a` (source vertex), `b` (edge), and `c` (destination vertex). The output automatically includes `b.srcid` and `b.destid` columns.

**Syntax**

```python Python theme={null}
map_triplets(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_edges_table,
    [ result_column_expressions [, ... ] ],
    [ result_edges_indexes [, ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). Used as the `a` and `c` vertices. |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). This function uses this argument as the `b` edge. |
| `result_schema` | str | A writable schema to create the result edges table. |
| `result_edges_table` | str | Name of the result edges table. The name must not conflict with the names of input tables. |
| `result_column_expressions` | List\[str] | Optional. SQL expressions that can reference `a` (source vertex), `b` (edge), and `c` (destination vertex). Use the `AS alias_name` keyword for stable names. |
| `result_edges_indexes` | List\[str] | Optional. Columns to index in the result edges table (e.g., `srcid` and `destid` columns). Specify an empty list for none. |

**Example**
Create a new triplet table from the vertices and edges tables with the `amount` and `same_country` columns, and generate indexes for the `src` and `destid` columns.

```python Python theme={null}
map_triplets(
    connection,
    "sales", "customers", "purchases",
    "sales", "purchases_triplets",
    [
      "b.amount AS amount",
      "CASE WHEN a.country = c.country THEN 1 ELSE 0 END AS same_country",
    ],
    ["srcid","destid"],
)
```

### reverse\_edges

Creates a new edges table with the `srcid` and `destid` columns reversed, preserving other columns. Use this function to traverse a graph in the opposite direction.

**Syntax**

```python Python theme={null}
reverse_edges(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_edges_table
    [ result_edges_indexes [, ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input edges table. |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the reversed edges table. |
| `result_edges_table` | str | Name of the reversed edges table to create. The name must not conflict with the names of input tables. |
| `result_edges_indexes` | List\[str] | Optional. Columns to index in the result edges table (e.g., `srcid` and `destid` columns). Specify an empty list for none. |

**Example**
Transform edge direction by reversing the `srcid` and `destid` columns. The example also creates indexes for these columns.

```python Python theme={null}
reverse_edges(
    connection,
    "sales", "purchases",
    "sales", "purchases_reversed",
    ["srcid","destid"],
)
```

### group\_edges

Groups duplicate rows of the `srcid` and `destid` columns, producing one row for each unique pair of values in a new edges table. This function performs aggregations based on one or more SQL expressions.

**Syntax**

```python Python theme={null}
group_edges(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_edges_table,
    [ result_column_expressions [, ... ] ],
    [ result_edges_indexes [, ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input edges table. |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the grouped tables. |
| `result_edges_table` | str | Name of the grouped edges table. The name must not conflict with the names of input tables. |
| `result_column_expressions` | List\[str] | One or more aggregate expressions for the `srcid` and `destid` columns. For details about SQL aggregations, see [Aggregate Functions](/aggregate-functions). <br /><br />Use the `AS alias_name` keyword for stable names. |
| `result_edges_indexes` | List\[str] | Optional. Columns to index in the result edges table (e.g., `srcid` and `destid` columns). Specify an empty list for none. |

**Example**
Create a new edge table that includes SQL aggregations for counting unique transactions `txn_count` and total sums `total_amount`. Also, this function generates indexes for the `src` and `destid` columns.

```python Python theme={null}
group_edges(
    connection,
    "sales", "purchases",
    "sales", "purchases_grouped",
    [
      "COUNT(*) AS txn_count",
      "SUM(amount) AS total_amount",
    ],
    ["srcid","destid"],
)
```

## Triplets

Produce triplet representations that are made of `a` (source vertex), `b` (edge), and `c` (destination vertex), either as a logical view or a materialized table for downstream queries.

### create\_triplets\_view

Creates a view that combines the edge table with the source and destination vertex attributes. This view is useful for analyzing relationships without having to repeatedly join tables. The view includes these columns:

* All original edge columns (including the `srcid` and `destid` columns).
* All source-vertex columns except `id`. Source-vertex column names have the `src_` prefix.
* All destination-vertex columns except `id`. Destination-vertex column names have the `dest_` prefix.

Use the [create\_triplets\_table](#create_triplets_table) function instead if you want to create a materialized table with indexes instead of a view.

**Syntax**

```python Python theme={null}
create_triplets_view(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_triplets_view
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input vertices and edges tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A schema to create the view. |
| `result_triplets_view` | str | Name of the triplets view to create. The name must not conflict with the names of input tables. |

**Example**
Create a triplets view to inspect edges with joined source and destination vertex attributes.

```python Python theme={null}
create_triplets_view(
    connection,
    "sales", "customers", "purchases",
    "sales", "triplets_v",
)
```

### create\_triplets\_table

Creates a materialized table that combines the edge table with the source and destination vertex attributes. This table is useful for analyzing relationships without having to repeatedly join tables. The created table includes these columns:

* All original edge columns (including the `srcid` and `destid` columns).
* All source-vertex columns except `id`. Source-vertex column names have the `src_` prefix.
* All destination-vertex columns except `id`. Destination-vertex column names have the `dest_` prefix.

Use the [create\_triplets\_view](#create_triplets_view) function if you want to create a view instead of a new table.

**Syntax**

```python Python theme={null}
create_triplets_table(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_triplets_table,
    [ result_triplets_indexes [, ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A schema to create the triplets table. |
| `result_triplets_table` | str | Name of the triplets table to create. The name must not conflict with any input names. |
| `result_triplets_indexes` | List\[str] | Optional. Columns to index in the result table. Specify an empty list for none. |

**Example**
Create a new table for triplets. Generate indexes for the `src_id` and `dest_id` columns.

```python Python theme={null}
create_triplets_table(
    connection,
    "sales", "customers", "purchases",
    "sales", "triplets_t",
    ["srcid","destid"],
)
```

## Degrees

Compute degree metrics for each vertex from the edges table. These functions produce small vertex tables suitable for joins and analytics.

### in\_degrees

Computes how many edges point to each vertex in an edge table by counting how many times each unique `destid` value appears. The result table has two columns: `id` (the destination vertex) and `in_degree` (the count).

**Syntax**

```python Python theme={null}
in_degrees(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_vertices_table,
    [ result_vertices_indexes [, ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input table. |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A schema to create the in-degree table. |
| `result_vertices_table` | str | Name of the vertices table to create (columns include `id` and  `in_degree`). |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Compute in-degrees per vertex and generate an index on the `id` column.

```python Python theme={null}
in_degrees(
    connection,
    "sales", "purchases",
    "sales", "customers_in_degree",
    ["id"],
)
```

### out\_degrees

Computes how many edges originate from each vertex in an edge table by counting how many times each unique `srcid` value appears. The result table has two columns: `id` (the source vertex) and `out_degree` (the count).

**Syntax**

```python Python theme={null}
out_degrees(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_vertices_table,
    [ result_vertices_indexes [, ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input table. |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A schema to create the out-degree table. |
| `result_vertices_table` | str | Name of the vertices table to create (columns include `id` and  `out_degree`). |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Compute the out-degrees count for each vertex and generate an index on the `id` column.

```python Python theme={null}
out_degrees(
    connection,
    "sales", "purchases",
    "sales", "customers_out_degree",
    ["id"],
)
```

### degrees

Computes the total degrees (in-degrees and out-degrees) for each vertex in an edge table by counting how many times each unique `srcid` and `destid` value appears. The result table has two columns: `id` (the destination or source vertex) and `degree` (the count).

**Syntax**

```python Python theme={null}
degrees(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_vertices_table,
    [ result_vertices_indexes [, ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input table. |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A schema to create the degree table. |
| `result_vertices_table` | str | Name of the vertices table to create (columns include `id` and  `out_degree`). |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Compute total degrees for each vertex and generate an index on the `id` column.

```python Python theme={null}
degrees(
    connection,
    "sales", "purchases",
    "sales", "customers_degree",
    ["id"],
)
```

## Vertex Extraction and Joins

Build vertex sets from edges and combine vertex attributes across tables. These functions are useful for shaping vertex properties and consolidating features.

### from\_edges

Builds a vertices table from an edges table by extracting the unique source and destination identifiers.

This function can optionally compute additional columns using SQL expressions by referencing the unique identifier as `ids.id`.

The created table always contains the `id` column with one additional column per expression.

**Syntax**

```python Python theme={null}
from_edges(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_vertices_table,
    [ result_column_expressions [ , ... ] ],
    [ result_vertices_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input table. |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A schema to create the result vertices table. |
| `result_vertices_table` | str | Name of the result vertices table (always includes `id`). |
| `result_column_expressions` | List\[str] | Optional. Specify a list of SQL expressions that reference `ids.id` to add additional columns. |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Create a vertices table from edge endpoints and add a `bucket` column that assigns each vertex to one of 10 buckets. Generate an index for the `id` and `bucket` columns.

```python Python theme={null}
from_edges(
    connection,
    "sales", "purchases",
    "sales", "customers_from_edges",
    ["ids.id % 10 AS bucket"],
    ["id","bucket"],
)
```

### join\_vertices

Merges two vertices tables by retaining every row from a primary table (`input_vertices_table`) and selectively updating rows that also appear in the modification table (`modification_vertices_table`). The merged table includes all vertices from the primary table that do not appear in the modification table.

For vertices that appear in both tables, the function must include a list of expressions (`resultAttributeExpressions`) in the same column order for every non-identifier column in the merged result table. These SQL expressions can add computations to columns, or simply add aliases if no changes are needed. Each expression can reference columns from the primary table (using alias `a`) or from the modification table (using alias `b`).

**Syntax**

```python Python theme={null}
join_vertices(
    connection,
    input_schema,
    input_vertices_table,
    modification_schema,
    modification_vertices_table,
    result_schema,
    result_vertices_table,
    result_attribute_expressions [ , ... ],
    [ result_vertices_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the primary input table. |
| `input_vertices_table` | str | Input vertices table (with the alias `a`). This table must have an `id` column. |
| `modification_schema` | str | A schema for the modification vertices table. |
| `modification_vertices_table` | str | Modification vertices table (with the alias `b`). This table must have an `id` column. |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create. |
| `result_attribute_expressions` | List\[str] | A list of SQL expressions that define the non-identifier columns of the joined result.<br /><br />Each expression can reference the left (`input_vertices_table`) vertex as `a` and the right (`modification_vertices_table`) vertex as `b`.<br /><br />You must end every expression with an explicit alias using `AS alias_name` so the output column names are stable. |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Merge vertex attributes and generate indexes for the `id` and `status` columns. This example includes two SQL expressions to update the `status` and `score` columns based on the modification vertex table using the `COALESCE` SQL reference function.

```python Python theme={null}
join_vertices(
    connection,
    "sales", "customers",
    "sales", "customers_updates",
    "sales", "customers_merged",
    [
      "COALESCE(b.new_status, a.status) AS status",
      "COALESCE(b.score_delta + a.score, a.score) AS score",
    ],
    ["id","status"],
)
```

### inner\_join\_vertices

Performs an inner join on two vertex tables using an equality comparison `a.id = b.id`. The result table automatically includes the `id` column from the first table.

The function must include a list of SQL expressions (`result_attribute_expressions`) in the same column order for every non-identifier column in the merged result table. These SQL expressions can add computations to columns, or simply add aliases if no changes are needed. Each expression can reference columns from the primary table using the alias `a` or from the modification table using the alias `b`.

**Syntax**

```python Python theme={null}
inner_join_vertices(
    connection,
    input_schema,
    input_vertices_table,
    other_schema,
    other_vertices_table,
    result_schema,
    result_vertices_table,
    result_attribute_expressions [ , ... ],
    [ result_vertices_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A schema for the left vertices table. |
| `input_vertices_table` | str | Input vertices table for the left side of the join (must have an `id` column). |
| `other_schema` | str | A schema for the right vertices table. |
| `other_vertices_table` | str | The other vertices table for the right side of the join (must have an `id` column). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create. |
| `result_attribute_expressions` | List\[str] | A list of SQL expressions that define the non-identifier columns of the joined result.<br /><br />Each expression can reference the left (`input_vertices_table`) vertex as `a` and the right (`other_vertices_table`) vertex as `b`.<br /><br />You must end every expression with an explicit alias using `AS alias_name` to ensure the output column names are stable. |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Create an inner join between two vertex tables and generate an index on the `id` column.

```python Python theme={null}
inner_join_vertices(
    connection,
    "sales", "customers",
    "sales", "profiles",
    "sales", "customers_joined",
    [
      "a.id AS id",
      "a.status AS status",
      "b.tier AS tier",
    ],
    ["id"],
)
```

### outer\_join\_vertices

Performs a left outer join between two vertices tables using an equality comparison `a.id = b.id`. The result table includes all rows from the left table. For left-table rows that have no match in the right table, any expression that reads columns from the right table with the alias `b` evaluates to NULL (while expressions that only read the table with the alias `a` remain non-NULL as usual).

The method must include a list of SQL expressions (`result_attribute_expressions`) in the same column order for every non-identifier column in the merged result table. These SQL expressions can add computations to columns, or simply add aliases if no changes are needed. Each expression can reference columns from the primary table using the alias `a` or from the modification table using the alias `b`.

**Syntax**

```python Python theme={null}
outer_join_vertices(
    connection,
    input_schema,
    input_vertices_table,
    other_schema,
    other_vertices_table,
    result_schema,
    result_vertices_table,
    result_attribute_expressions [ , ... ],
    [ result_vertices_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A schema for the left vertices table. |
| `input_vertices_table` | str | Left vertices table (must have an `id` column). |
| `other_schema` | str | A schema for the right vertices table (with the alias `b`). |
| `other_vertices_table` | str | Right vertices table (must have an `id` column). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create. |
| `result_attribute_expressions` | List\[str] | A list of SQL expressions that define the non-identifier columns of the joined result.<br /><br />Each expression can reference the left (`input_vertices_table`) vertex as `a` and the right (`other_vertices_table`) vertex as `b`.<br /><br />You must end every expression with an explicit alias using `AS alias_name` to ensure the output column names are stable. |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Perform a left outer join on two vertices tables and generate an index on the `id` column.

```python Python theme={null}
outer_join_vertices(
    connection,
    "sales", "customers",
    "sales", "profiles",
    "sales", "customers_joined",
    [
      "a.id AS id",
      "a.status AS status",
      "b.tier AS tier",
    ],
    ["id"],
)
```

### collect\_neighbors

For each vertex in a table, this function collects information on neighbors (identifier and any attributes) as an array of tuples. For a specified direction (`IN`, `OUT`, or `BOTH`), the function aggregates tuples representing each neighboring vertex into an array.

The direction types are:

* `IN` — Neighbors with edges pointing to the vertex (edges where `destid = id`).
* `OUT` — Neighbors that the vertex points to (edges where `srcid = id`).
* `BOTH` — Union of `IN` and `OUT` with neighbors from incoming (`destid = id`) and outgoing (`srcid = id`) edges.

The result table has the columns `id` (the vertex identifier) and `neighbors` (an array of tuples representing each neighbor).

If an error occurs after table creation, the function drops the result table.

**Syntax**

```python Python theme={null}
collect_neighbors(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_table,
    direction,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the neighbors collection table (with the `id` and `neighbors` columns). |
| `direction` | EdgeDirection | Specifies the direction of traversal. Supported values are:<br />`IN` — Neighbors that have edges pointing to the vertex (edges where `destid = id`).<br />For example, if an edge `5` points to `10`, then for `id=10`, neighbor `5` is included. <br /><br />`OUT` — Neighbors that the vertex points to (edges where `srcid = id`). <br />For example: If an edge `5` points to `10`, then for `id=5`, neighbor `10` is included.<br /><br />`BOTH` — The union of `IN` and `OUT`. This traversal includes neighbors from edges pointing to `id` and edges from `id`. |
| `result_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`).  Specify an empty list for none. |

**Example**
Collect incoming neighbors for each vertex and generate an index on the `id` column. The `direction` argument set to `IN` collects neighbors pointing to `id`.

```python Python theme={null}
collect_neighbors(
    connection,
    "sales", "customers", "purchases",
    "sales", "neighbors_in",
    EdgeDirection.IN,
    ["id"],
)
```

### collect\_edges

For each vertex in a table, this function collects an array of adjacent edge rows based on the specified direction. The result table has two columns: `id` (the vertex identifier) and `edges` (an array of tuples, each tuple containing all columns from the edges table for a connected edge).

The `direction` types are:

* `IN` — Edges pointing to the vertex (edges where `destid = id`).
* `OUT` — Edges originating from the vertex (edges where `srcid = id`).
* `BOTH` — Union of `IN` and `OUT` that includes edges from incoming (`destid = id`) and outgoing (`srcid = id`) directions. This direction retains duplicates.

If an error occurs after table creation, the function drops the result table.

**Syntax**

```python Python theme={null}
collect_edges(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_table,
    direction,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the edge collection table (with the `id` and `edges` columns). |
| `direction` | EdgeDirection | Specifies the direction of traversal. Supported values are:<br />`IN` — Neighbors that have edges pointing to the vertex (edges where `destid = id`).<br />For example, if an edge `5` points to `10`, then for `id=10`, neighbor `5` is included. <br /><br />`OUT` — Neighbors that the vertex points to (edges where `srcid = id`). For example, if an edge `5` points to `10`, then for `id=5`, neighbor `10` is included.<br /><br />`BOTH` — The union of `IN` and `OUT`. This traversal includes neighbors from edges pointing to `id` and edges from `id`. |
| `result_indexes` | List\[str] | Optional. Columns to index. Specify an empty list for none. |

**Example**
Collect outgoing edges per vertex. The example sets the `direction` to `OUT` to collect edges from `id`.

```python Python theme={null}
collect_edges(
    connection,
    "sales", "customers", "purchases",
    "sales", "outgoing_edges",
    EdgeDirection.OUT,
    ["id"],
)
```

## Algorithms

High-level graph algorithms that iterate over the graph structure to produce labels, components, or counts.

### label\_propagation

Executes the [Label Propagation Algorithm](https://en.wikipedia.org/wiki/Label_propagation_algorithm) (LPA) to assign community labels to vertices.

Each vertex starts with its own identifier as its label. For a number set by the `maxIterations` argument, each vertex updates its label to the most frequent label among its neighbors. The algorithm determines ties by choosing the smallest label. The algorithm uses temporary tables for intermediate results and drops these tables when the process completes or if it fails. Isolated vertices retain their initial label. The final table stores `id` and `label` columns and can include indexes.

**Syntax**

```python Python theme={null}
label_propagation(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_vertices_table,
    max_iterations,
    result_vertices_indexes,
    [ result_vertices_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create (columns include `id` and `label`). |
| `max_iterations` | int | Maximum number of iterations (must be `1` or greater). |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Run label propagation for 10 iterations and assign labels to vertices. Generate an index on the `id` column.

```python Python theme={null}
label_propagation(
    connection,
    "sales", "customers", "purchases",
    "sales", "lpa_labels",
    10,
    ["id"],
)
```

<h3 id="connected_components">
  connected\_components
</h3>

Identifies the connected components of an undirected graph. This algorithm configures a Pregel computation in which each vertex initially sets its component label equal to its own identifier `id`.

In each iteration, vertices send their component label to neighbors. Each vertex updates based on the aggregated minimum value of its current component label and any received values. The process repeats until no more updates occur.

The result table maps each vertex `id` to its final `component` label.

**Syntax**

```python Python theme={null}
connected_components(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_vertices_table,
    [ result_vertices_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create. |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**

Compute connected components and generate an index on the `id` column.

```python Python theme={null}
connected_components(
    connection,
    "sales", "customers", "purchases",
    "sales", "components",
    ["id"],
)
```

### strongly\_connected\_components

Computes strongly connected components (SCC) in a directed graph. This function runs a recursive algorithm that partitions vertices into subsets where every vertex is reachable from other vertices in the same subset.

This function uses recursive partitioning. The algorithm selects a pivot (typically the minimum identifier `id`), computes its predecessor set (vertices that can reach the pivot), and its descendant sets (vertices reachable from the pivot). Then, the function identifies the SCC as their intersection, removes that SCC from the graph, and recurses on the remainder until all vertices have been assigned to an SCC. The output contains columns for the `id` and `component` identifiers (the minimum `id` in the SCC).

The function creates temporary tables in the result schema to store intermediate results. This function drops these tables when the computation completes or fails. The final result table contains two columns: `id` (vertex identifier) and `component` (the minimum vertex identifier in its SCC subset).

**Syntax**

```python Python theme={null}
strongly_connected_components(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_vertices_table,
    [ result_vertices_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create (columns include `id` and `component`). |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Compute the SCC and generate an index on the `id` column.

```python Python theme={null}
strongly_connected_components(
    connection,
    "sales", "customers", "purchases",
    "sales", "scc",
    ["id"],
)
```

### TriangleCount

TriangleCount identifies all 3-cycles (triangles) in the graph and counts how many distinct triangles each vertex participates in.

The algorithm first builds a canonical, undirected edge set by ensuring `srcid < destid` and removing duplicates to prevent double-counting. If your input edges are already canonicalized and deduplicated, use `TriangleCount.run_pre_canonicalized` to skip preprocessing for faster performance.

The function then counts triangles (`a`, `b`, `c`) where `a < b < c` by intersecting neighbor lists and aggregates per-vertex participation to produce a result table with the `id` and `triangle_count` columns.

`run` **syntax**

```python Python theme={null}
TriangleCount.run(
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_vertices_table,
    [ result_vertices_indexes [ , ... ] ],
)
```

`run_pre_canonicalized` **syntax**

```python Python theme={null}
TriangleCount.run_pre_canonicalized(
    connection,
    input_schema,
    input_vertices_table,
    canonical_edges_schema,
    canonical_edges_table,
    result_schema,
    result_vertices_table,
    [ result_vertices_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns).<br /><br />For `runPreCanonicalized `execution, the function assumes this table is already canonicalized (e.g., `srcid < destid`) and deduplicated. |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create (columns include `id` and `triangle_count`). |
| `result_vertices_indexes` | list\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Examples**

**Count Triangles Using** `run`

Canonicalize the raw edges internally, count unique triangles, and write per-vertex triangle counts with an index on the `id` column.

```python Python theme={null}
TriangleCount.run(
    connection,
    "sales",
    "customers",
    "purchases",
    "sales",
    "triangle_counts",
    ["id"],
)
```

**Count Triangles Using** `run_pre_canonicalized`

Use a pre-canonicalized, deduplicated edge table to count triangles and write per-vertex triangle counts with an index on the `id` column.

```python Python theme={null}
TriangleCount.run_pre_canonicalized(
    connection,
    "sales",
    "customers",
    "sales",
    "purchases_canonical",
    "sales",
    "triangle_counts",
    ["id"],
)
```

### pregel

Provides a generic vertex‑centered iteration framework for custom graph algorithms, similar to the [Pregel model](https://research.google/pubs/pregel-a-system-for-large-scale-graph-processing/).

Each iteration updates vertex states by sending messages along edges and then aggregating these messages to compute new states. The algorithm continues iterating until it reaches convergence (no state changes or no messages produced) or a specified iteration cap.

The algorithm uses multiple specified SQL expressions.

**Syntax**

```python Python theme={null}
pregel(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_vertices_table,
    initializer_expr,
    send_to_source_expr,
    send_to_dest_expr,
    aggregate_expr,
    updater_expr,
    max_iterations,
    [ result_vertices_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create (columns include `id` and `result`). |
| `initializer_expr` | str | A SQL expression to compute the initial state for each vertex (e.g., `CASE WHEN type = 'seed' THEN 1.0 ELSE 0.0 END`). |
| `send_to_source_expr` | str | Optional. Defines the message sent to the source vertex of an edge, referencing the current vertex states `a.state` (source) and `c.state` (destination), and any edge attributes in `b`.<br />Example: `CASE WHEN a.country = c.country THEN 1 ELSE 0 END` |
| `send_to_dest_expr` | str | Optional. Defines the message sent to the destination vertex of an edge, referencing the current vertex states `a.state` (source) and `c.state` (destination), and any edge attributes as `b.<edge_column>`. <br />Example: `CASE WHEN a.country = c.country THEN 1 ELSE 0 END` |
| `aggregate_expr` | str | A SQL aggregation to combine messages per vertex (e.g., `SUM(msg)`).<br />For a list of supported aggregations, see [Aggregate Functions](/aggregate-functions). |
| `updater_expr` | str | A SQL expression to compute the next state from the current state and aggregated messages. <br /><br />For example, this code updates the state to the minimum current state and the aggregated message (similar to the [connected\_components](#connected_components) function): `LEAST(a.state, COALESCE(m.aggregated_message, a.state)) AS state` |
| `max_iterations` | int | Maximum number of iterations (must be either `-1` or a value of `1` or greater). <br /><br />If this value is `-1`, the Pregel algorithm has no limit, and it runs until convergence. <br /><br />A value of `0` causes the algorithm to return an error. |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |

**Example**
Run a simple Pregel computation summing incoming edge amounts into the vertex state for 10 iterations at most, and generate an index on the `id` column.

```python Python theme={null}
pregel(
    connection,
    "sales", "customers", "purchases",
    "sales", "pregel_result",
    "0 AS state",
    "b.amount",
    None,
    "SUM(msg) AS aggregated_message",
    "state + COALESCE(aggregated_message, 0) AS state",
    10,
    ["id"],
)
```

## Paths & Ranking

These functions include the shortest-path and PageRank algorithms.

### shortest\_paths

Computes the shortest distance from every vertex to each set of landmark vertices using an iterative relaxation algorithm. The algorithm resembles [Bellman–Ford](https://en.wikipedia.org/wiki/Bellman%E2%80%93Ford_algorithm) but simultaneously handles multiple destinations.

Each landmark starts at distance `0` and all others at positive infinity. On each iteration, the algorithm examines every edge and checks whether traveling through the connected neighbor would yield a shorter route to a landmark. If a shorter route exists, the algorithm updates the distance of the source vertex. The process stops when no distances improve or the algorithm reaches the maximum number of iterations.

After the process finishes, the algorithm writes a result table with the `srcid`, `destid`, and `distance` columns.

**Syntax**

```python Python theme={null}
shortest_paths(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_table,
    landmarks[ , ... ],
    edge_weight_column,
    max_iterations,
    [ result_vertices_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result table. |
| `result_table` | str | Name of the distances result table. This table has the `srcid`, `destid`, and `distance` columns. |
| `landmarks` | List\[int] | One or more landmark vertex identifiers. This list must be non-empty and contain no NULLs. |
| `edge_weight_column` | Optional\[str] | Optional. The name of an edge weight column. <br />If you do not specify this argument, the default value is `1.0`. |
| `max_iterations` | int | Maximum number of relaxation iterations. This value must be `1` or greater. |
| `result_indexes` | List\[str] | Optional. The list of columns to index in the result table (`srcid`, `destid`, and `distance` columns). <br /><br />Specify an empty list for none. |

**Example**
Compute distances from landmarks and generate indexes on the `src` and `dest` columns.

```python Python theme={null}
shortest_paths(
    connection,
    "sales", "customers", "purchases",
    "sales", "distances",
    [1, 42],
    None,
    10,
    ["srcid","destid"],
)
```

### static\_page\_rank

Computes [PageRank](https://en.wikipedia.org/wiki/PageRank) scores for each vertex over a fixed number of iterations.

The algorithm follows the standard PageRank formula with a damping factor (`damping_factor`) and uses common table expressions to calculate contributions from incoming edges and redistribute ranks from dangling nodes.

The algorithm supports two variants:

* Standard PageRank — All vertices start with rank `1.0/N`, where `N` is the number of vertices. Specify this variant if `personalizationSrcId` is `null`.
* Personalized PageRank — The specified vertex starts with a rank of `1.0`, while others start with a rank of `0.0`. Specify this variant if `personalizationSrcId` is a vertex identifier.

After running PageRank for a fixed number of iterations, the function writes a result vertices table containing all original vertex columns with a new PageRank scoring column.

**Syntax**

```python Python theme={null}
static_page_rank(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_vertices_table,
    num_iterations,
    damping_factor,
    [ result_vertices_indexes [ , ... ] ],
    personalization_src_id,
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create (columns include all vertex columns and a `pagerank` column). |
| `num_iterations` | int | Number of iterations to run. This value must be 1 or greater. |
| `damping_factor` | float | The damping factor, a value between `0.0` and `1.0`, that controls how the PageRank random surfer moves across vertices. <br /><br />A high value (e.g., `0.85`) puts more emphasis on link structure, encouraging the algorithm to move from each vertex to a neighbor. This value means PageRank scores tend to concentrate around well-linked regions.<br /><br />A lower value (e.g., `0.50`) allows the algorithm to ignore edges and instead jump to a vertex chosen from a base distribution. For standard PageRank, this base is uniform over all vertices. For personalized PageRank, the base is biased toward the specified vertex. This behavior creates more uniform scoring with less sensitivity to link topology.<br /><br />If this value is outside the range `0.0` to `1.0`, the function raises an error. |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |
| `personalization_src_id` | Optional\[int] | Determines whether PageRank uses the standard or the personalized variant.  <br /><br />For personalized mode, specify a vertex identifier. PageRank scoring starts with this vertex set at `1.0`, while all other vertices start at `0.0`. <br /><br />For standard mode, specify `None`.  All vertices start with rank `1.0/N`, where `N` is the number of vertices. |

**Example**
Run fixed-iteration PageRank and generate an index on the `id` column. This example uses a damping factor of `0.85` to ensure the ranking concentrates on highly linked regions.

```python Python theme={null}
static_page_rank(
    connection,
    "sales", "customers", "purchases",
    "sales", "pagerank_static",
    10,
    0.85,
    ["id"],
    None,
)
```

### dynamic\_page\_rank

Computes PageRank scores until convergence based on a specified threshold value (`tolerance`). Unlike the [static\_page\_rank](#static_page_rank) function, this algorithm runs iterations until the sum of absolute differences between ranks in successive iterations is less than or equal to the `tolerance` value. The algorithm handles personalization similarly to `static_page_rank`. At each iteration, the function uses the PageRank formula, collects rank values, and redistributes them.

The algorithm supports two variants:

* Standard PageRank — All vertices start with rank `1.0/N`, where `N` is the number of vertices. Specify this variant if `personalizationSrcId` is `null`.
* Personalized PageRank — The specified vertex starts with a rank of `1.0`, while others start with a rank of `0.0`. Specify this variant if `personalizationSrcId` is a vertex identifier.

After running PageRank until the system reaches the `tolerance` threshold, the function writes a vertices table containing all the original vertex columns with a new PageRank scoring column.

**Syntax**

```python Python theme={null}
dynamic_page_rank(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_vertices_table,
    tolerance,
    damping_factor,
    [ result_vertices_indexes [ , ... ] ],
    personalization_src_id,
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create (columns include all vertex columns and `pagerank`). |
| `tolerance` | float | The convergence threshold that stops iterations when the rank changes become negligible. <br /><br />After each iteration, the algorithm measures the maximum change in any rank of a vertex. If that change is below the specified tolerance, the algorithm considers the computation converged and stops early. <br /><br />A high tolerance value (e.g., `0.001`) is good for quick exploratory runs that generate an approximate ranking. <br /><br />A low tolerance value (e.g., `0.000001`) generates a precise ranking, but requires a higher compute cost. |
| `damping_factor` | float | The damping factor, a value between `0.0` and `1.0`, that controls how the PageRank random surfer moves across vertices. <br /><br />A high value (e.g., `0.85`) puts more emphasis on link structure, encouraging the algorithm to move from each vertex to a neighbor. The high value means PageRank scores tend to concentrate around well-linked regions.<br /><br />A lower value (e.g., `0.50`) allows the algorithm to ignore edges and instead jump to a vertex chosen from a base distribution. For standard PageRank, this base is uniform over all vertices. For personalized PageRank, the base is biased toward the specified vertex. This behavior creates more uniform scoring with less sensitivity to link topology.<br /><br />If this value is outside the range `0.0` to `1.0`, the function raises an error. |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |
| `personalization_src_id` | Optional\[int] | Determines whether PageRank uses the standard or the personalized variant.  <br /><br />For personalized mode, specify a vertex identifier. PageRank scoring starts with this vertex set at `1.0`, while all other vertices start at `0.0`. <br /><br />For standard mode, specify `None`. All vertices start with rank `1.0/N`, where `N` is the number of vertices. |

**Example**
Run dynamic PageRank to convergence and generate an index on the `id` column. This example uses a low `tolerance` value of `1.0e-6`, which generates high-precision rankings but requires more computing resources.

```python Python theme={null}
dynamic_page_rank(
    connection,
    "sales", "customers", "purchases",
    "sales", "pagerank_dynamic",
    1.0e-6,
    0.85,
    ["id"],
    None,
)
```

<Note>
  You can specify the deprecated `reset_prob` keyword argument instead of the `damping_factor` argument. The value means the same thing (the damping factor, and not the teleport probability). Specifying the `reset_prob` argument raises the `DeprecationWarning` warning. You cannot specify both `damping_factor` and `reset_prob` arguments in the same function call, but you must specify one of them.
</Note>

### stable\_marriage

Computes a stable matching between two groups using the [Gale–Shapley algorithm](https://en.wikipedia.org/wiki/Gale%E2%80%93Shapley_algorithm). Each suitor proposes to candidates in order of preference, and each candidate tentatively holds the best proposal received. The algorithm iterates until no suitor can improve or until it reaches `max_iterations` rounds, producing a suitor-optimal stable matching.

The suitor preferences table must have the `suitor_id`, `candidate_id`, and `rank` columns. The candidate preferences table must have the `candidate_id`, `suitor_id`, and `rank` columns. All `rank` values must be non-NULL. The result table contains the `suitor_id` and `candidate_id` columns.

**Syntax**

```python Python theme={null}
stable_marriage(
    connection,
    input_schema,
    input_suitor_prefs_table,
    input_candidate_prefs_table,
    result_schema,
    result_table,
    max_iterations,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_suitor_prefs_table` | str | Suitor preferences table (must have `suitor_id`, `candidate_id`, and `rank` columns). |
| `input_candidate_prefs_table` | str | Candidate preferences table (must have `candidate_id`, `suitor_id`, and `rank` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `suitor_id` and `candidate_id` columns). |
| `max_iterations` | int | Maximum number of proposal rounds that must be `1` or greater. |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `suitor_id`). Specify an empty list for none. |

**Example**

Compute a stable matching between suitors and candidates with a maximum of 20 rounds.

```python Python theme={null}
stable_marriage(
    connection,
    "sales", "suitor_prefs", "candidate_prefs",
    "sales", "matches",
    20,
    ["suitor_id"],
)
```

### jaccard\_similarity

Computes the [Jaccard similarity](https://en.wikipedia.org/wiki/Jaccard_index) for every pair of vertices that share at least one neighbor. The Jaccard similarity between two vertices is the size of the intersection of their neighbor sets divided by the size of their union. This metric is useful for link prediction and duplicate detection.

The result table contains the `srcid`, `destid`, and `similarity` columns.

**Syntax**

```python Python theme={null}
jaccard_similarity(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_table,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `srcid`, `destid`, and `similarity` columns). |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `srcid` and `destid`). Specify an empty list for none. |

**Example**

Compute pairwise Jaccard similarity scores.

```python Python theme={null}
jaccard_similarity(
    connection,
    "sales", "purchases",
    "sales", "jaccard_scores",
    ["srcid", "destid"],
)
```

### cosine\_similarity

Computes the [cosine similarity](https://en.wikipedia.org/wiki/Cosine_similarity) for every pair of vertices that share at least one neighbor. The cosine similarity represents the cosine of the angle between two neighbor-set vectors. This metric is useful for recommendation systems and measuring structural equivalence.

The result table contains the `srcid`, `destid`, and `similarity` columns.

**Syntax**

```python Python theme={null}
cosine_similarity(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_table,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `srcid`, `destid`, and `similarity` columns). |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `srcid` and `destid`). Specify an empty list for none. |

**Example**

Compute pairwise cosine similarity scores.

```python Python theme={null}
cosine_similarity(
    connection,
    "sales", "purchases",
    "sales", "cosine_scores",
    ["srcid", "destid"],
)
```

### k\_core\_decomposition

Computes the [k-core decomposition](https://en.wikipedia.org/wiki/Degeneracy_\(graph_theory\)#k-Cores) of a graph by iteratively removing vertices with degree less than `k` until only the maximal subgraph with the minimum degree `k` remains. The algorithm computes the coreness value for each vertex.

The result table contains the `id` and `core` columns.

**Syntax**

```python Python theme={null}
k_core_decomposition(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_table,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `id` and `core` columns). |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `id`). Specify an empty list for none. |

**Example**

Compute k-core decomposition and index on the `id` column.

```python Python theme={null}
k_core_decomposition(
    connection,
    "sales", "customers", "purchases",
    "sales", "kcore_result",
    ["id"],
)
```

### eigenvector\_centrality

Computes [eigenvector centrality](https://en.wikipedia.org/wiki/Eigenvector_centrality) scores for each vertex by iteratively updating the score of each vertex to be the sum of the scores of its neighbors, followed by normalization. Vertices connected to other high-scoring vertices receive higher centrality.

The result table contains the `id` and `centrality` columns.

**Syntax**

```python Python theme={null}
eigenvector_centrality(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_table,
    num_iterations,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `id` and `centrality` columns). |
| `num_iterations` | int | Number of power-iteration rounds that must be `1` or greater. |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `id`). Specify an empty list for none. |

**Example**

Compute eigenvector centrality with 20 iterations.

```python Python theme={null}
eigenvector_centrality(
    connection,
    "sales", "customers", "purchases",
    "sales", "eigenvec_result",
    20,
    ["id"],
)
```

### cycle\_detection

Detects vertices that participate in cycles in a directed graph using iterative degree-based peeling. The algorithm repeatedly removes vertices with an in-degree or out-degree of zero until no such vertices remain. The remaining vertices are those involved in at least one cycle.

The result table includes the `id` column for each vertex in a cycle.

**Syntax**

```python Python theme={null}
cycle_detection(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    result_schema,
    result_table,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `id` column). |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `id`). Specify an empty list for none. |

**Example**

Detect vertices involved in cycles.

```python Python theme={null}
cycle_detection(
    connection,
    "sales", "customers", "purchases",
    "sales", "cycle_vertices",
    ["id"],
)
```

### max\_bipartite\_matching

Computes [maximum matching](https://en.wikipedia.org/wiki/Maximum_matching) in a bipartite graph using an iterative augmenting-path approach. The algorithm greedily matches unmatched vertices and then refines the matching until it finds no further augmenting path or until it reaches the number of rounds as specified by the `max_iterations` value. The result is a set of edges where no two edges share a vertex.

The result table contains the `srcid` and `destid` columns.

**Syntax**

```python Python theme={null}
max_bipartite_matching(
    connection,
    input_schema,
    input_edges_table,
    result_schema,
    result_table,
    max_iterations,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). The graph must be bipartite, where the source and destination vertices must form two disjoint sets. |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `srcid` and `destid` columns). |
| `max_iterations` | int | Maximum number of augmentation rounds that must be `1` or greater. |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `srcid` and `destid`). Specify an empty list for none. |

**Example**

Compute a maximum bipartite matching with up to 10 augmentation rounds.

```python Python theme={null}
max_bipartite_matching(
    connection,
    "sales", "purchases",
    "sales", "bipartite_matches",
    10,
    ["srcid", "destid"],
)
```

### louvain\_modularity\_optimization

Detects communities using the [Louvain method](https://en.wikipedia.org/wiki/Louvain_method), a greedy modularity optimization algorithm. The algorithm iteratively moves vertices between communities to maximize modularity, then aggregates communities and repeats until modularity stops improving or until it reaches the number of rounds as specified by the `max_iterations` value.

The result table contains the `id` and `community` columns.

**Syntax**

```python Python theme={null}
louvain_modularity_optimization(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    weight_col,
    result_schema,
    result_table,
    max_iterations,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `weight_col` | str or None | The name of the edge weight column. Specify `None` for unweighted graphs. The default weight is `1.0`. |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `id` and `community` columns). |
| `max_iterations` | int | Maximum number of iterations that must be `1` or greater. |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `id`). Specify an empty list for none. |

**Example**

Detect communities using the Louvain method with a maximum of 10 iterations on an unweighted graph.

```python Python theme={null}
louvain_modularity_optimization(
    connection,
    "sales", "customers", "purchases",
    None,
    "sales", "louvain_communities",
    10,
    ["id"],
)
```

### aggregate\_messages

A general-purpose message-passing primitive that sends messages along edges and aggregates them at destination vertices. This function is useful for building custom graph computations, feature engineering, and implementing algorithms not available as built-in functions.

For each edge, the function evaluates a message expression referencing source vertex attributes (`a.*`), edge attributes (`b.*`), and destination vertex attributes (`c.*`). The function then aggregates the messages for each destination vertex using a SQL aggregation function.

The result table contains the vertex `id` and the aggregated message column, named using the `AS` alias in the `aggregate_expr` argument.

**Syntax**

```python Python theme={null}
aggregate_messages(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    message_expr,
    aggregate_expr,
    result_schema,
    result_table,
    [ result_indexes [ , ... ] ]
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `message_expr` | str | A SQL expression defining the message to send along each edge. The expression can reference source vertex columns as `a.*`, edge columns as `b.*`, and destination vertex columns as `c.*`. Example: `"b.amount"` |
| `aggregate_expr` | str | A SQL aggregation expression that combines the messages arriving at each vertex. Reference the value produced by the `message_expr` argument as `msg`, and name the result column with an `AS` alias. For a list of supported aggregations, see [Aggregate Functions](/aggregate-functions). Example: `"SUM(msg) AS total"` |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create. |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `id`). Specify an empty list for none. |

**Example**

Compute the sum of incoming edge weights for each vertex.

```python Python theme={null}
aggregate_messages(
    connection,
    "sales", "customers", "purchases",
    "b.amount",
    "SUM(msg) AS total_amount",
    "sales", "vertex_totals",
    ["id"],
)
```

### a\_star\_shortest\_path

Computes the shortest path between a source and target vertex using the [A\* search algorithm](https://en.wikipedia.org/wiki/A*_search_algorithm). A\* extends Dijkstra's algorithm with a specified heuristic expression that estimates the remaining cost to the target, guiding the search toward the goal, reducing the number of vertices explored, and stopping early if it reaches the number of iterations as specified by the `max_iterations` value.

The heuristic must be admissible, meaning it never overestimates the true cost, for the algorithm to guarantee an optimal path. The result table contains the `srcid`, `destid`, `cost`, and `path_index` columns, which represent the ordered nodes on the shortest path.

**Syntax**

```python Python theme={null}
a_star_shortest_path(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    source_node_id,
    target_node_id,
    weight_col,
    heuristic_expr,
    result_schema,
    result_table,
    [ result_indexes [ , ... ] ],
    max_iterations,
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `source_node_id` | int | The vertex identifier of the source node. |
| `target_node_id` | int | The vertex identifier of the target node. |
| `weight_col` | str | The name of the edge weight column. Specify `None` for unweighted graphs. The default weight is `1.0`. |
| `heuristic_expr` | str | A SQL expression that estimates the remaining cost from a vertex to the target. Reference vertex columns using the `v` alias, and substitute literal values for the target coordinates. For example, `"ABS(v.x - 10) + ABS(v.y - 20)"` is the Manhattan distance heuristic for the target at `(10, 20)`. |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create. |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `srcid` and `destid`). Specify an empty list for none. |
| `max_iterations` | int | Optional. Maximum number of relaxation iterations that must be `1` or greater. Default: `10000`. |

**Example**

Find the shortest path from vertex `1` to vertex `42` using edge weights and a coordinate-based heuristic.

```python Python theme={null}
a_star_shortest_path(
    connection,
    "sales", "customers", "purchases",
    1, 42,
    "weight",
    "ABS(v.x - 10) + ABS(v.y - 20)",
    "sales", "astar_path",
    ["srcid", "destid"],
    100,
)
```

### yens\_k\_shortest\_paths

Computes the `k` shortest loopless paths between a source and target vertex using [Yen's algorithm](https://en.wikipedia.org/wiki/Yen%27s_algorithm). The algorithm iteratively finds the next shortest path by deviating from previously discovered paths at each spur node, using up to the number of Dijkstra iterations as specified by the `max_iterations` value for each spur-node search.

The result table contains the `path_id`, `srcid`, `destid`, `cost`, and `path_index` columns, which represent the ordered nodes on each path.

**Syntax**

```python Python theme={null}
yens_k_shortest_paths(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    source_node_id,
    target_node_id,
    k,
    weight_col,
    result_schema,
    result_table,
    [ result_indexes [ , ... ] ],
    max_iterations,
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `source_node_id` | int | The vertex identifier of the source node. |
| `target_node_id` | int | The vertex identifier of the target node. |
| `k` | int | The number of shortest paths to find. This value must be `1` or greater. |
| `weight_col` | str | The name of the edge weight column. Specify `None` for unweighted graphs. The default weight is `1.0`. |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create. |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `path_id` and `srcid`). Specify an empty list for none. |
| `max_iterations` | int | Optional. Maximum number of Dijkstra iterations per sub-path that must be `1` or greater. Default: `10000`. |

**Example**

Find the three shortest paths from vertex `1` to vertex `42` using edge weights.

```python Python theme={null}
yens_k_shortest_paths(
    connection,
    "sales", "customers", "purchases",
    1, 42,
    3,
    "weight",
    "sales", "yen_paths",
    ["path_id", "srcid"],
    100,
)
```

### leiden\_modularity\_optimization

Detects communities using the [Leiden algorithm](https://en.wikipedia.org/wiki/Leiden_algorithm), a refinement of the Louvain method that guarantees the formation of well-connected communities. The algorithm optimizes modularity through iterative local moves and a refinement phase that prevents poorly connected communities, repeating until modularity stops improving or until it reaches the number of rounds as specified by the `max_iterations` value.

The result table contains the `id` and `community` columns.

**Syntax**

```python Python theme={null}
leiden_modularity_optimization(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    weight_col,
    resolution_param,
    result_schema,
    result_table,
    max_iterations,
    [ result_indexes [ , ... ] ],
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `weight_col` | str | The name of the edge weight column. Specify `None` for unweighted graphs. The default weight is `1.0`. |
| `resolution_param` | float | Resolution parameter controlling the granularity of communities. Higher values produce more communities. A typical value is `1.0`. |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `id` and `community` columns). |
| `max_iterations` | int | Maximum number of iterations that must be `1` or greater. |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `id`). Specify an empty list for none. |

**Example**

Detect communities using the Leiden algorithm with a resolution of `1.0` and a maximum of 10 iterations.

```python Python theme={null}
leiden_modularity_optimization(
    connection,
    "sales", "customers", "purchases",
    "weight",
    1.0,
    "sales", "leiden_communities",
    10,
    ["id"],
)
```

### time\_respecting\_shortest\_path

Computes the shortest path between a source and target vertex while respecting the temporal ordering of edges. Each edge has a timestamp, and the algorithm traverses only edges whose timestamps are nondecreasing along the path, modeling real-world scenarios where events must occur in chronological order.

The timestamp column must be of `TIMESTAMP` type. The result table contains the `srcid`, `destid`, `cost`, and `path_index` columns.

**Syntax**

```python Python theme={null}
time_respecting_shortest_path(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    source_node_id,
    target_node_id,
    time_col,
    weight_col,
    earliest_departure,
    max_wait_time,
    result_schema,
    result_table,
    [ result_indexes [ , ... ] ],
    max_iterations,
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `source_node_id` | int | The vertex identifier of the source node. |
| `target_node_id` | int | The vertex identifier of the target node. |
| `time_col` | str | The name of the `TIMESTAMP` column on edges representing the event time. |
| `weight_col` | str | The name of the edge weight column. Specify `None` for unweighted graphs. The default weight is `1.0`. |
| `earliest_departure` | str | A `TIMESTAMP` literal specifying the earliest allowed departure time from the source. Example: `"2024-01-01 00:00:00"` |
| `max_wait_time` | Optional\[int] | Optional. Maximum waiting time in seconds at any vertex. Specify `None` to wait indefinitely. |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create. |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `srcid` and `destid`). Specify an empty list for none. |
| `max_iterations` | int | Optional. Maximum number of relaxation iterations that must be `1` or greater. Default: `10000`. |

**Example**

Find the time-respecting shortest path from vertex `1` to vertex `42` departing no earlier than midnight on January 1, 2024.

```python Python theme={null}
time_respecting_shortest_path(
    connection,
    "sales", "customers", "purchases",
    1, 42,
    "event_time",
    "weight",
    "2024-01-01 00:00:00",
    None,
    "sales", "trsp_path",
    ["srcid", "destid"],
    100,
)
```

### temporal\_page\_rank

Computes [PageRank](https://en.wikipedia.org/wiki/PageRank) scores using exponential time decay on edges. Recent edges contribute more to the ranking than older edges, making this algorithm suitable for graphs where recency matters (e.g., communication networks, transaction logs).

The time column must be of `TIMESTAMP` type. The algorithm supports both standard and personalized variants, similar to [static\_page\_rank](#static_page_rank).

**Syntax**

```python Python theme={null}
temporal_page_rank(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    time_col,
    decay_factor,
    reference_time,
    result_schema,
    result_vertices_table,
    num_iterations,
    damping_factor,
    [ result_vertices_indexes [ , ... ] ],
    personalization_src_id,
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `time_col` | str | The name of the `TIMESTAMP` column on edges. |
| `decay_factor` | float | Controls the rate of exponential decay. A higher value causes older edges to lose influence faster. |
| `reference_time` | str | A `TIMESTAMP` literal specifying the reference point for computing time differences. Example: `"2024-12-31 23:59:59"` |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_vertices_table` | str | Name of the vertices table to create (columns include all vertex columns and a `pagerank` column). |
| `num_iterations` | int | Number of iterations to run that must be `1` or greater. |
| `damping_factor` | float | The damping factor, specified as a value between `0.0` and `1.0`. For details, see [static\_page\_rank](#static_page_rank). |
| `result_vertices_indexes` | List\[str] | Optional. Columns to index in the result vertices table (e.g., `id`). Specify an empty list for none. |
| `personalization_src_id` | Optional\[int] | For the personalized mode, specify a vertex identifier. For the standard mode, specify `None`. |

**Example**

Run temporal PageRank with a decay factor of `0.5` using a reference time of December 31, 2024.

```python Python theme={null}
temporal_page_rank(
    connection,
    "sales", "customers", "purchases",
    "event_time",
    0.5,
    "2024-12-31 23:59:59",
    "sales", "temporal_pr",
    10,
    0.85,
    ["id"],
    None,
)
```

### random\_walk

Generates configurable random walks from each vertex using [Node2Vec](https://arxiv.org/abs/1607.00653)-style biased sampling. The return parameter (`p_return` argument) and in-out parameter (`q_inout` argument) control whether the walk favors revisiting the previous node (breadth-first) or exploring further (depth-first).

The result table contains the `walk_id`, `step`, and `node_id` columns.

**Syntax**

```python Python theme={null}
random_walk(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    walk_length,
    num_walks_per_node,
    p_return,
    q_inout,
    result_schema,
    result_table,
    [ result_indexes [ , ... ] ],
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `walk_length` | int | Length of each random walk (number of steps). |
| `num_walks_per_node` | int | Number of walks to generate from each vertex. |
| `p_return` | float | Return parameter, where a lower value makes the walk more likely to return to the previous node (breadth-first behavior). |
| `q_inout` | float | In-out parameter, where a lower value makes the walk more likely to explore further from the previous node (depth-first behavior). |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `walk_id`, `step`, and `node_id` columns). |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `walk_id`). Specify an empty list for none. |

**Example**

Generate five random walks of length 10 from each vertex using Node2Vec-style parameters.

```python Python theme={null}
random_walk(
    connection,
    "sales", "customers", "purchases",
    10,
    5,
    1.0,
    1.0,
    "sales", "walks",
    ["walk_id"],
)
```

### fast\_random\_projection

Generates node embeddings using the [Fast Random Projection](https://arxiv.org/abs/2005.02140) (FastRP) algorithm. FastRP creates low-dimensional vector representations of vertices by iteratively averaging neighbor embeddings with sparse random projections. These embeddings are useful as features for downstream machine learning tasks.

The result table contains the `id` and `embedding` columns. The `embedding` column is a vector with length equal to the `embedding_dim` value.

**Syntax**

```python Python theme={null}
fast_random_projection(
    connection,
    input_schema,
    input_vertices_table,
    input_edges_table,
    embedding_dim,
    iterations,
    normalize,
    result_schema,
    result_table,
    [ result_indexes [ , ... ] ],
)
```

| **Argument** | **Data Type** | **Description** |
| - | - | - |
| `connection` | pyocient.Connection | An active database connection using the `pyocient` module. |
| `input_schema` | str | A non-empty schema containing the input tables. |
| `input_vertices_table` | str | Input vertices table (must have an `id` column). |
| `input_edges_table` | str | Input edges table (must have the `srcid` and `destid` columns). |
| `embedding_dim` | int | Dimensionality of the output embedding vectors. |
| `iterations` | int | Number of neighbor-averaging iterations. Higher values capture more distant structural information. |
| `normalize` | bool | If you set this argument to `True`, the function scales each embedding vector to unit length using the L2 norm. |
| `result_schema` | str | A writable schema to create the result tables. |
| `result_table` | str | Name of the result table to create (includes the `id` and `embedding` columns). |
| `result_indexes` | List\[str] | Optional. Columns to index in the result table (e.g., `id`). Specify an empty list for none. |

**Example**

Generate 64-dimensional normalized embeddings with three iterations.

```python Python theme={null}
fast_random_projection(
    connection,
    "sales", "customers", "purchases",
    64,
    3,
    True,
    "sales", "embeddings",
    ["id"],
)
```

## Bibliography

Brin, Sergey, and Lawrence Page. "The Anatomy of a Large-Scale Hypertextual Web Search Engine." *Computer Networks and ISDN Systems* 30, no. 1–7 (1998): 107–117. [https://en.wikipedia.org/wiki/PageRank](https://en.wikipedia.org/wiki/PageRank).

Blondel, Vincent D., Jean-Loup Guillaume, Renaud Lambiotte, and Etienne Lefebvre. "Fast Unfolding of Communities in Large Networks." *Journal of Statistical Mechanics: Theory and Experiment* 2008, no. 10 (2008): P10008. [https://en.wikipedia.org/wiki/Louvain\_method](https://en.wikipedia.org/wiki/Louvain_method).

Chen, Haochen, Syed Fahad Sultan, Yingtao Tian, Muhao Chen, and Steven Skiena. "Fast and Accurate Network Embeddings via Very Sparse Random Projection." *Proceedings of the 28th ACM International Conference on Information and Knowledge Management* (2019). [https://arxiv.org/abs/1908.11512](https://arxiv.org/abs/1908.11512).

Gale, David, and Lloyd S. Shapley. "College Admissions and the Stability of Marriage." *The American Mathematical Monthly* 69, no. 1 (1962): 9–15. [https://en.wikipedia.org/wiki/Gale–Shapley\_algorithm](https://en.wikipedia.org/wiki/Gale%E2%80%93Shapley_algorithm).

Grover, Aditya, and Jure Leskovec. "node2vec: Scalable Feature Learning for Networks." *Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining* (2016): 855–864. [https://arxiv.org/abs/1607.00653](https://arxiv.org/abs/1607.00653).

Hart, Peter E., Nils J. Nilsson, and Bertram Raphael. "A Formal Basis for the Heuristic Determination of Minimum Cost Paths." *{IEEE} Transactions on Systems Science and Cybernetics* 4, no. 2 (1968): 100–107. [https://en.wikipedia.org/wiki/A\*\_search\_algorithm](https://en.wikipedia.org/wiki/A*_search_algorithm).

Jaccard, Paul. "Étude comparative de la distribution florale dans une portion des Alpes et des Jura." *Bulletin de la Société Vaudoise des Sciences Naturelles* 37 (1901): 547–579. [https://en.wikipedia.org/wiki/Jaccard\_index](https://en.wikipedia.org/wiki/Jaccard_index).

Malewicz, Grzegorz, Matthew H. Austern, Aart J. C. Bik, James C. Dehnert, Ilan Horn, Naty Leiser, and Grzegorz Czajkowski. "Pregel: A System for Large-Scale Graph Processing." *Proceedings of the 2010 ACM SIGMOD International Conference on Management of Data* (2010): 135–146. [https://research.google/pubs/pregel-a-system-for-large-scale-graph-processing/](https://research.google/pubs/pregel-a-system-for-large-scale-graph-processing/).

Seidman, Stephen B. "Network Structure and Minimum Degree." *Social Networks* 5, no. 3 (1983): 269–287. [https://en.wikipedia.org/wiki/Degeneracy\_(graph\_theory)#k-Cores](https://en.wikipedia.org/wiki/Degeneracy_\(graph_theory\)#k-Cores).

Traag, Vincent A., Ludo Waltman, and Nees Jan van Eck. "From Louvain to Leiden: Guaranteeing Well-Connected Communities." *Scientific Reports* 9, no. 1 (2019): 5233. [https://en.wikipedia.org/wiki/Leiden\_algorithm](https://en.wikipedia.org/wiki/Leiden_algorithm).

Yen, Jin Y. "Finding the K Shortest Loopless Paths in a Network." *Management Science* 17, no. 11 (1971): 712–716. [https://en.wikipedia.org/wiki/Yen%27s\_algorithm](https://en.wikipedia.org/wiki/Yen%27s_algorithm).

"Cosine Similarity." Wikipedia. Accessed August 2026. [https://en.wikipedia.org/wiki/Cosine\_similarity](https://en.wikipedia.org/wiki/Cosine_similarity).

"Eigenvector Centrality." Wikipedia. Accessed August 2026. [https://en.wikipedia.org/wiki/Eigenvector\_centrality](https://en.wikipedia.org/wiki/Eigenvector_centrality).

"Maximum Matching." Wikipedia. Accessed August 2026. [https://en.wikipedia.org/wiki/Maximum\_matching](https://en.wikipedia.org/wiki/Maximum_matching).

## Related Links

[OCGraph Java Library](/ocgraph-java-library)

[Ocient Python Module: pyocient](/ocient-python-module-pyocient)
