This list of options contains model options for all models in the Ocient® System.
Association Rules
Model Options
Optional
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
temporalOrdering — If you set this option to true, the model treats antecedent items as those appearing earlier in time and consequent items as those appearing later, within the same transaction. Also, you must specify the transactionColumn and timeColumn options. The training input is a per-item table with shape (transaction_id, time, item) rather than the default unordered ARRAY per row. The default value is false, and the model preserves the unordered behavior unchanged.
timeColumn — The name of the column in the training SELECT SQL statement that orders items within a transaction. This option is required when you set the temporalOrdering option to true. Otherwise, the database throws a syntax error. The model ignores this option when you do not set the temporalOrdering option or set it to false. The column must be a numeric or temporal type that supports <, >, and MAX() expressions.
transactionColumn — The name of the column in the training SELECT SQL statement that groups items into one basket or sequence. This option is required when you set the temporalOrdering option to true. Otherwise, the database throws a syntax error. The model ignores this option when you do not set the temporalOrdering option or set it to false.
Backward Feature Elimination
Model Options
Required
evaluationMetric — This option specifies the criterion that drives the search loop. For classification inner models, supported values are accuracy, f1_macro, f1_weighted, mcc, auc_roc, and log_loss. For regression inner models, supported values are r_squared, rmse, mae, and mape.
innerModelType — This option specifies the model type for the inner supervised learner that the wrapper model repeatedly trains on candidate column subsets. The value must be a model type that supports classification or regression. Examples include logistic regression, multiple linear regression, and decision tree.
Optional
holdoutFraction — If you set this option, the value must be a positive double in the range (0, 0.5] that represents the fraction of the training data the wrapper model holds out for trial scoring. The model trains trials on the complementary partition and scores on the holdout partition. The default value is 0.2.
innerModelOptions — If you set this option, the value must be a JSON object where the keys are option names that are accepted by the inner model type. The wrapper model forwards these options verbatim to every CREATE MLMODEL SQL statement the model executes. The default value is an empty object.
maxFeatures — This option specifies the maximum number of features. The Forward Feature Selection model honors this value. Backward Feature Elimination accepts this value but does not use it. The default value is unlimited.
maxIterations — If you set this option, the value must be a positive integer that represents the maximum number of search rounds. The default value is 100.
metricDirection — If you set this option, the value must match the natural direction of the chosen evaluationMetric option: higher_is_better for accuracy, f1_macro, f1_weighted, mcc, auc_roc, and r_squared; lower_is_better for log_loss, rmse, mae, and mape. This option acts as a sanity-check assertion; it cannot override the metric’s natural direction. The default is the natural direction of the chosen evaluationMetric option value.
metricThreshold — If you set this option, the value is an absolute score threshold at which the search terminates early. This option is mutually exclusive with the relativeAcceptable option.
minFeatures — This option specifies the minimum number of features. The Backward Feature Elimination model honors this value as the floor cardinality of the surviving subset. The Forward Feature Selection model accepts this value but does not use it. The default value is 1.
parallelism — If you set this option, the value must be a positive integer that bounds how many candidate trial trainings the wrapper model runs concurrently within a single round. The default value is 5. The model limits actual concurrency to the range [1, min(numCandidates, hardware_concurrency())].
relativeAcceptable — If you set this option, the value is a percentage that represents how far below the full-feature reference score the wrapper model accepts. This option is mutually exclusive with the metricThreshold option.
Bagging
Model Options
Required
baseModels — This option specifies the children of the bagging model. You must specify this value as a JSON array, where each object in the array has the three fields type, count, and options.
taskType — This option specifies the type of task, which must be either CLASSIFICATION or REGRESSION depending on the type of model for training.
Optional
ROCNumSamples — If you set this option, you must specify a positive integer that represents the number of samples for the model to use when calculating the area under the ROC curve. You must also set the metrics option to true. The default value is the number of child models.
bootstrap — If you set this option to true, the model uses bootstrap sampling with replacement, meaning each child model trains on a random subset of the data (either the rowsPerChild or fractionSelected value sets the exact number of rows), and the same row can appear multiple times for each child. If you set this option to false, the model does not use replacement, meaning each row can appear at most one time per child. The default value is false.
continuousFeatures — If you set this option, the value must be a comma-separated list of the feature indexes that are continuous numeric variables. Indexes start with 1. In the default state, the model considers no features as continuous.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
fractionSelected — If you set this option, the option represents the proportion of rows the model uses to train each child model. The value is a double that must be in the interval (0, 1]. You cannot set this option if you also set the rowsPerChild option to a positive value. The default behavior is that the model uses all available rows.
inputsPerChild — If you set this option, the option represents the number of features used to create each child model. The default value is the total number of features divided by 3 and rounded up.
maxChildThreads — If you set this option, the value must be an integer representing the maximum number of threads each child model can use. If a child accepts a maxThreads option, the model passes this value to the child.
maxThreads — If you set this option, the option represents the maximum number of parallel threads to use while the model trains. This value must be a positive integer. The default value is 16.
metrics — If you set this option to true, the system calculates certain metrics depending on the value of the taskType option. If you set the taskType option to CLASSIFICATION, the metrics are the percentage of correctly classified rows and the area under the ROC curve. If you set the value to REGRESSION, the metrics are the root mean square error and the adjusted R-squared. The default value is false.
requiredFeatures — If you set this option, the value must be a comma-separated list of integers representing features starting at index 1. The bagging model passes these features down to every child. The default value is an empty list, meaning there is no required feature.
rowsPerChild — If you set this option to a positive integer, the number represents the number of rows (from a random sample) to use for each decision tree. If you set this option to 0, each child uses all available rows. The default value is 0. You cannot set this option to a positive value if you also set the fractionSelected option.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
weighted — If you set this option, the system passes the value directly to child models that support it. The behavior depends on the model type of the target child.
Boosting
Model Options
Required
baseModels — This option specifies the children of the bagging model. You must specify this as a JSON array, where each object in the array has the three fields type, count, and options.
learningRate — A decimal value between 0.0 and 1.0 that tunes how much the model learns from each successive child.
taskType — This option specifies the type of task, which must be either CLASSIFICATION or REGRESSION depending on the type of model for training.
Optional
ROCNumSamples — If you set this option, you must specify a positive integer that represents the number of samples for the model to use when calculating the area under the ROC curve. You must also set the metrics option to true. The default value is 10.
continuousFeatures — If you set this option, the value must be a comma-separated list of the feature indexes that are continuous numeric variables. Indexes start with 1. In the default state, the model considers no features as continuous.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
fractionSelected — If you set this option, the option represents the proportion of rows the model uses to train each child model. The value is a double that must be in the interval (0, 1]. You cannot set this option if you also set the rowsPerChild option to a positive value. The default behavior is that the model uses all available rows.
inputsPerChild — If you set this option, the value must be an integer type that is greater than or equal to 1, which specifies the number of input features each boosting child should use. This value cannot exceed the number of input features available in the data set. When you specify this value, the algorithm deterministically cycles through pre-enumerated feature subsets to ensure each child uses exactly the specified number of features. When you do not specify this value, the model uses all available features for each child.
lossFunction — If you set this option, the value represents the loss function used by the model. Accepted values are: 'squared_error' and 'log_loss'. When you set this value to 'squared_error', the model calculates errors as the squared difference between predicted and actual values. The target column must contain numeric values. This is the default value when the taskType option is set to REGRESSION. When you set this value to 'log_loss', the model calculates errors using logistic loss. This is the default value when the taskType option is CLASSIFICATION.
maxThreads — If you set this option, the option represents the maximum number of parallel threads to use while the model trains. This value must be a positive integer. The default value is 16.
metrics — If you set this option to true, the system calculates certain metrics depending on the value of the taskType option. If the taskType option is set to CLASSIFICATION, the metrics are the percentage of correctly classified rows and the area under the ROC curve. If the taskType option is set to REGRESSION, the metrics are the root mean square error and the adjusted R-squared. The default value is false.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
Dbscan
Model Options
Required
epsilon — The distance threshold for the eps-neighborhood the system uses during computation of the k-nearest neighbors (KNN) graph. You must set this option to a finite positive double.
minPts — The minimum number of neighbors (including the row itself) required for a row to qualify as a core point. You must set this value to a positive integer.
Optional
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
insertBatchSize — If you set this option, the value must be a positive integer that starts at 1. This value controls how many (row_id, cluster_id) pairs the model batches into each INSERT INTO temp.<...> VALUES (...) statement when materializing the cores table at the end of training. Smaller batches under-utilize the per-statement INSERT overhead; larger batches stress the SQL parser. The default value is 131072 (= 128 * 1024).
kNeighbors — If you set this option, the value must be a positive integer that is at least as large as the minPts option value. This value controls the number of nearest neighbors the system uses when materializing the KNN graph at training time. Higher values trade training cost for better border-point recall at the fixed epsilon option value. The default value is max(20, minPts * 4).
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
metric — If you set this option, the value must be one of 'euclidean', 'cosine', or 'manhattan'. This value selects the distance metric for the KNN graph and inference. The default value is 'euclidean'.
normalize — If you set this option to true, the database normalizes each feature column to the zero mean and unit standard deviation before computing distances. The default value is true.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
Decision Tree
Model Options
Optional
ROCNumSamples — If you set this option, you must specify a positive integer that represents the number of samples for the model to use when calculating the area under the ROC curve. You must also set the metrics option to true. The default value is 10.
continuousFeatureBatchSize — If you set this option, the value must be a positive integer. This value sets the maximum number of continuous features where the system combines split-aggregate queries into a single UNION ALL query during decision tree training. Higher values reduce per-query dispatch and compile overhead, especially for wide-feature datasets, but increase peak SQL Node memory holding the combined result set. The worst-case combined result set scales with this value multiplied by the numClasses and numSplits values. The default value is 256, which matches the distinctCountLimit value of the discrete path so that per-chunk peak result-set memory stays comparable to (and in practice well below) what the discrete path already produces in a single UNION ALL query across all discrete features without chunking. Lower this value if you observe SQL Node memory pressure on workloads with very high numClasses or numSplits values; raise it for narrow workloads with low cardinality where dispatch overhead dominates. This option only affects continuous features. The discrete path already uses a single UNION ALL across all features per node.
continuousFeatures — If you set this option, the value must be a comma-separated list of the feature indexes that are continuous numeric variables. Indexes start with 1. In the default state, the model considers no features as continuous.
distinctCountLimit — If you set this option, the value must be a positive integer type. This value sets the limit for how many distinct values a non-continuous feature and label can contain. The default value is 256.
doPrune — If you set this option to true, the model uses Pessimistic Error Pruning (PEP) to prune the tree after training. The default value is false.
enableResplits — If you set this option, it must be a boolean type that determines if the tree can reuse the same continuous feature multiple times along a single branch (e.g., split on x1 < 7 and later x1 < 3). This action can capture more complex, range-specific relationships. The default value is true, meaning that continuous features remain available for additional splits after use, thereby allowing the tree to create more complex decision boundaries. If you set this option to false, the model marks continuous features as exhausted after their first use, and the model cannot use them again in subsequent splits in the same tree.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
featureSubsetStrategy — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the option specifies how many features the decision tree should consider at each split from the still-available features. When this value is higher, the model has a higher accuracy and lower variance, but takes longer to train. You can specify this option either as an integer (e.g., 4, meaning consider up to four features at each split) or one of these values: all (check every feature), sqrt (check up to the square root of the number of total features), log2 (check up to the base-2 logarithm of the number of total features), and one-third (check up to one-third of the number of total features). The default value is all.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
maxCellsToFetch — If you set this option, the value must be a positive integer. Controls the chunking behavior when fetching feature values during model training. The limit represents the maximum number of data cells (calculated as number of columns × number of rows) that can be fetched in a single operation, not a byte limit. When the expected data size exceeds this threshold, the algorithm switches to database-based processing using SQL queries instead of in-memory processing. This value defaults to 33,554,432 cells (calculated as 32 × 1024 × 1024).
maxDepth — If you set this option, the value must be a positive integer. This value sets the maximum allowable depth of the decision tree (the maximum number of features to split on). The default is unspecified, which means there is no maximum depth.
maxLeafNodes — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be an integer greater than or equal to 2. This value caps the number of leaf nodes in the decision tree during the standard depth-first build. When the tree reaches this limit, the model stops queuing further splits and converts pending nodes to leaves; this is a soft cap rather than the best-first growth, so the resulting tree might differ from a best-first tree of the same leaf budget. The default value is unspecified, which means there is no limit on the number of leaf nodes.
maxRows — If you set this option, the value must be a positive integer. This option limits the number of rows used for model training by creating a snapshot table with only the specified number of rows from the input query. When this option is unspecified, the model trains using all rows from the input query.
maxThreads — If you set this option, the value must be a positive integer. This value indicates the maximum number of parallel threads to use while the model trains. The default value is 2.
metrics — If you set this option to true, the model also calculates the percentage of samples correctly classified by the model and saves this information in a catalog table. This option defaults to false.
minImpurityDecrease — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a non-negative double. The tree splits only if the decrease in impurity (the difference between the parent node impurity and the weighted sum of child node impurities) is greater than or equal to this value. The default value is 0.0, meaning the model accepts all splits that improve impurity.
minSamplesLeaf — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer. This value sets the minimum number of samples required in a leaf node. The model rejects any split that creates a child node with fewer samples than this value, making the node a leaf instead. The default value is 1, meaning every leaf must have at least one sample.
minSamplesSplit — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be an integer greater than or equal to 2. This value sets the minimum number of samples required to split a node. Nodes with fewer samples than this value become leaves. The default value is 2.
numSplits — If you set this option, the value must be an integer greater than 1. This value sets the maximum number of binary branches a continuous feature can consider. The default value is 32.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
skipLimitCheck — If you set this option to true, the model skips cardinality checks that throw errors when columns have too many values. The limit that this option checks is the same one specified by the distinctCountLimit option. The default value is false.
splitMetric — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the option controls which function the model uses to evaluate the quality of a split during tree construction. Supported options are: gini_impurity (measures impurity based on class distributions), entropy (measures impurity using information entropy, i.e., -SUM(p_k * ln(p_k))), and hellinger (uses the Hellinger distance, which requires exactly two output classes). The default value is gini_impurity.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
suppressJIT — If you set this option to true, the model suppresses just-in-time code generation.
weighted — If you set this option, the model considers weights for labels. If you set this option value to true, you must specify an additional column as a double in the training data for label weights. Rows with the same labels must have the same weights. If you set this value to auto, the model calculates weights automatically by weighting each label according to the ratio of the count of the most frequent label to the count of the specified label. As a result, the most frequent label has the weight 1.0 and the other label weights are higher. The default value is false, which means all labels have equal weight.
Forward Feature Selection
Model Options
Required
evaluationMetric — This option specifies the criterion that drives the search loop. For classification inner models, supported values are accuracy, f1_macro, f1_weighted, mcc, auc_roc, and log_loss. For regression inner models, supported values are r_squared, rmse, mae, and mape.
innerModelType — This option specifies the model type for the inner supervised learner that the wrapper model repeatedly trains on candidate column subsets. The value must be a model type that supports classification or regression. Examples include logistic regression, multiple linear regression, and decision tree.
Optional
holdoutFraction — If you set this option, the value must be a positive double in the range (0, 0.5] that represents the fraction of the training data the wrapper model holds out for trial scoring. The model trains trials on the complementary partition and scores on the holdout partition. The default value is 0.2.
innerModelOptions — If you set this option, the value must be a JSON object where the keys are option names that are accepted by the inner model type. The wrapper model forwards these options verbatim to every CREATE MLMODEL SQL statement the model executes. The default value is an empty object.
maxFeatures — This option specifies the maximum number of features. The Forward Feature Selection model honors this value. Backward Feature Elimination accepts this value but does not use it. The default value is unlimited.
maxIterations — If you set this option, the value must be a positive integer that represents the maximum number of search rounds. The default value is 100.
metricDirection — If you set this option, the value must match the natural direction of the chosen evaluationMetric option: higher_is_better for accuracy, f1_macro, f1_weighted, mcc, auc_roc, and r_squared; lower_is_better for log_loss, rmse, mae, and mape. This option acts as a sanity-check assertion; it cannot override the metric’s natural direction. The default is the natural direction of the chosen evaluationMetric option value.
metricThreshold — If you set this option, the value is an absolute score threshold at which the search terminates early. This option is mutually exclusive with the relativeAcceptable option.
minFeatures — This option specifies the minimum number of features. The Backward Feature Elimination model honors this value as the floor cardinality of the surviving subset. The Forward Feature Selection model accepts this value but does not use it. The default value is 1.
parallelism — If you set this option, the value must be a positive integer that bounds how many candidate trial trainings the wrapper model runs concurrently within a single round. The default value is 5. The model limits actual concurrency to the range [1, min(numCandidates, hardware_concurrency())].
relativeAcceptable — If you set this option, the value is a percentage that represents how far below the full-feature reference score the wrapper model accepts. This option is mutually exclusive with the metricThreshold option.
Gaussian Discriminant Analysis
Model Options
Optional
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
metrics — If you set this option to true, the model calculates classification quality metrics (accuracy, MCC, AUC ROC, and per-class precision, recall, and F1) on the validation data (or the training data if you do not specify the VALIDATE ON clause) and stores them in the sys.gaussian_discriminant_analysis_models system catalog table. The default value is false.
normalize — If you set this option to true, the model automatically computes the mean and standard deviation of each feature. Then, the model normalizes the training data using the mean and standard deviation and stores these values so that the inference normalizes inputs identically. The default value is true.
priors — This option specifies how the model computes class priors P(c). If you set this option to 'empirical', the model uses observed class frequencies in the training data. If you set this option to 'uniform', the model assigns each of the K classes a prior of 1/K, which is the correct choice when the class proportions in the training data do not reflect the population. The default value is 'empirical'.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
regParam — If you set this option, the value must be a non-negative double representing a ridge term added to each per-class covariance diagonal before inversion (Sigma_c <- Sigma_c + regParam * I). Use a strictly positive value when classes have low row counts or when strongly correlated features render the unregularized covariance rank-deficient. The default value is 0.0.
suppressArrayLengthCheck — If you set this option, you must set the featureArray option to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
Gaussian Mixture Model
Model Options
Required
numDistributions — This option must be a positive integer type that specifies the number of clusters of Gaussian distributions for the model to make.
Optional
covarianceType — (BETA. This option is part of an in-development feature and is subject to change.) Specifies the type of covariance matrix to use. The values are: full (each component has its own full covariance matrix), diag (each component has its own diagonal covariance matrix), spherical (each component has a single variance shared across all features), or tied (all components share the same full covariance matrix). The default value is full.
epsilon — If you specify this option, the value must be a valid positive floating point number. When the maximum distance that the entire best model moves in its n-dimensional space is less than this value, the algorithm terminates. The default value is 0.00000001 (1e-8).
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
maxIterations — If you set this option, the option represents the maximum number of optimization iterations to train the model. For higher values of this option, the model is likelier to converge to the expected epsilon, but it might take longer to train. The default value is 100.
normalize — If you set this option to true, the model automatically computes the mean and standard deviation of each feature and uses them to normalize the data during training. Defaults to true.
numInitializations — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value specifies the number of times the model runs the full training procedure with different random k-means initializations while keeping the result with the best log-likelihood. The default value is 1.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
Gradient Boosted Trees
Model Options
Required
learningRate — A decimal value between 0.0 and 1.0 that tunes how much the model learns from each successive child.
numChildren — An integer value representing the total number of trees to build sequentially. Each tree learns to correct the errors of the previous trees.
Optional
ROCNumSamples — If you set this option, you must specify a positive integer that represents the number of samples for the model to use when calculating the area under the ROC curve. You must also set the metrics option to true. The default value is 10.
continuousFeatures — If you set this option, the value must be a comma-separated list of the feature indexes that are continuous numeric variables. Indexes start with 1. In the default state, the model considers no features as continuous.
earlyStoppingRounds — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a non-negative integer. Training stops if the loss does not improve for this number of consecutive trees. The model computes the loss on the training data (the same data used to fit each tree). The value 0 disables early stopping. The default value is 0.
earlyStoppingTolerance — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a non-negative double. This value sets the minimum relative improvement in loss required to count as an improvement for early stopping. A new loss L counts as an improvement only if L < bestLoss - tolerance * |bestLoss|. Use this option only when the earlyStoppingRounds value is greater than 0. The default value is 0.0001 (1e-4).
enableResplits — If you set this option, the value must be a boolean type that determines if the tree can reuse the same continuous feature multiple times along a single branch (e.g., split on x1 < 7 and later x1 < 3). This action can capture more complex, range-specific relationships. The default value is true, meaning that continuous features remain available for additional splits after use, thereby allowing the tree to create more complex decision boundaries. If you set this option to false, the model marks continuous features as exhausted after their first use, and the model cannot use them again in subsequent splits in the same tree.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
featureSubsetStrategy — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the model passes this option directly to the child regression trees. The option specifies how many features each child tree should consider at each split from the still-available features. When this value is higher, the model has higher accuracy and lower variance, but takes longer to train. You can specify this option either as an integer (e.g., 4, meaning consider up to 4 features at each split) or one of these values: all (checks every feature), sqrt (checks up to the square root of the number of total features), log2 (checks up to the base-2 logarithm of the number of total features), and one-third (checks up to one-third of the number of total features). The default value is all.
fractionSelected — If you set this option, the option represents the proportion of rows the model uses to train each child model. The value is a double that must be in the interval (0, 1]. You cannot set this option if you also set the rowsPerChild option to a positive value. The default behavior is that the model uses all available rows.
inputsPerChild — If you set this option, the value must be an integer type greater than or equal to 1 that specifies the number of input features each boosting tree should use. This value cannot exceed the number of input features available in the data set. When you specify this value, the algorithm deterministically cycles through pre-enumerated feature subsets to ensure each tree uses exactly the specified number of features. The default behavior is that the model uses all available features for each tree.
lossFunction — If you set this option, the option represents the loss function used and determines the type of task the model does. Accepted values are: 'squared_error' and 'log_loss'. When you set this value to 'squared_error', the model calculates errors as the squared difference between predicted and actual values. The target column must contain numeric values. This is the default value for regression tasks. When you set this value to 'log_loss', the model calculates errors using logistic loss. This is the default value for classification tasks.
maxCellsToFetch — If you set this value, the value must be an integer type that determines the memory threshold to switch from training with system memory to training with SQL queries in the database. In-memory training is generally faster, but is limited by the available SQL Node memory. If the size of a training data subset exceeds this value, then the system performs training operations using SQL queries. The default value is 33,554,432 (calculated as 32 * 1024 * 1024).
maxDepth — If you set this value, the value must be a positive integer type that represents the maximum allowable depth of the child trees. The default value is 3.
maxLeafNodes — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be an integer greater than or equal to 2. The model passes this option directly to the child regression trees. This value caps the number of leaf nodes in each child tree during its depth-first build (it is a soft cap rather than best-first growth). The default value is unspecified, which means there is no limit on the number of leaf nodes.
maxThreads — If you set this value, the value must be a positive integer type that sets the maximum number of parallel threads to use for training each child decision tree. Parallel threads do not affect the sequential method of training each tree. The default value is 16.
metrics — If you set this value to true, the system calculates and stores final model metrics (R²/RMSE for regression or Accuracy/LogLoss for classification) on the training data. The default value is false.
minImpurityDecrease — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a non-negative double. The child regression tree splits only if the decrease in impurity is greater than or equal to this value. The model passes this option directly to the child regression trees. The default value is 0.0.
minSamplesLeaf — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer. This value sets the minimum number of samples required in a leaf node for each child regression tree. The model passes this option directly to the child regression trees. The default value is 1.
minSamplesSplit — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be an integer greater than or equal to 2. This value sets the minimum number of samples required to split a node for each child regression tree. The model passes this option directly to the child regression trees. The default value is 2.
numSplits — If you set this option, the value must be an integer greater than 1. This value sets the maximum number of binary branches a continuous feature can consider. The default value is 32.
regAlpha — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a non-negative double. This value sets the L1 regularization strength applied as a post-hoc soft-thresholding of the leaf prediction t of each child tree: the model replaces t with sign(t) * max(0, |t| - regAlpha). This formula is a leaf-output transform that the model applies after it builds the child regression tree. The formula is not the L1 term in the second-order objective used by XGBoost, so values do not directly transfer between the two systems. The default value is 0.0 (no L1 regularization).
regLambda — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a non-negative double. This value sets the L2 regularization strength applied as a post-hoc shrinkage of the leaf prediction of each child tree: the model multiplies the leaf output by 1 / (1 + regLambda). This formula is a leaf-output transform that the model applies after it builds the child regression tree. The formula is not the L2 term in the closed-form leaf-weight objective used by XGBoost, so values do not directly transfer between the two systems. The default value is 0.0 (no L2 regularization).
resplitDepth — If you set this option, the value must be an integer type that sets the maximum depth at which tree nodes can be re-split during optimization. This option controls how deep the algorithm searches for better split points. The default value is 6.
resplitThreshold — If you set this option, the value must be a decimal type that sets the minimum improvement threshold required to trigger a re-split operation. Lower values allow more aggressive re-splitting but can increase training time. Higher values require larger improvements to trigger re-splits. The default value is 0.1.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
taskType — If you set this option, the value determines the task the model performs. Accepted values are 'REGRESSION' (regression) and 'CLASSIFICATION' (classification). When you set this option to 'REGRESSION', the model sets the default loss function to 'squared_error'. When you set this option to 'CLASSIFICATION', the model sets the default loss function to 'log_loss'. You can also set the lossFunction option explicitly, but it must be compatible with the specified task type. If you set neither the taskType nor lossFunction options, the model defaults to regression with the squared_error loss function.
Kmeans
Model Options
Required
k — This option must be a positive integer type that specifies how many clusters to make.
Optional
epsilon — (BETA. This option is part of an in-development feature and is subject to change.) If you specify this option, the value must be a valid positive floating point value. When the maximum distance that a centroid moves from one iteration of the algorithm to the next is less than this value, the algorithm terminates. The default value is 0.0001 (1e-4).
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
lloydRounds — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the option represents the maximum number of iterations of the Lloyd algorithm to train the model after guessing the centroids. For higher values of this option, the model is more likely to be accurate but takes longer to train. The value 0 indicates an unlimited number of iterations, in which case the algorithm runs only until convergence (see the epsilon option). The default value is 20.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
normalize — If you set this option to true, the model normalizes the data before the start of training. The default value is true.
numInitializations — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value specifies the number of times the model runs the full k-means algorithm with different random initializations and keeps the result with the lowest total within-cluster sum of squares (inertia). Higher values increase the chance of finding the globally optimal clustering, but increase training time proportionally. The default value is 1.
oversampling — If you set this option, the option represents the number of candidate guesses for the model to choose in the parallel-round phase of k-means||. For higher values of this option, the model is more likely to be accurate but takes longer to train. The default value is k.
parallelRounds — If you set this option, the option represents the minimum number of parallel rounds for which the k-means|| algorithm runs. For higher values of this option, the model is more likely to be accurate but takes longer to train. The default value is 8.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
K Nearest Neighbors
Model Options
Required
k — This option must be a positive integer type that specifies how many closest points to use for classifying a new point.
Optional
distance — If you set this option, the value must be a function in SQL syntax for calculating the distance between a point used for classification and points in the training data set. This function should use the variables x1, x2, … for the 1st, 2nd, … features in the training data set, and p1, p2, … for the features in the point for classification. The default value is the Euclidean distance function.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
normalize — If you set this option to true, the model automatically computes the mean and standard deviation of each feature and uses them to normalize the data during training. The default value is true.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
weight — If you specify this option, the value must be a function in SQL syntax for calculating the weight of a neighbor. The function should use the variable d for distance. By default, the distance is set to 1.0 / (d + 0.1), thus avoiding division by zero on exact inputs and still allowing neighbors to have some influence.
Linear Combination Regression
Model Options
Required
functionN — You must specify the first function using a key named 'function1'. Subsequent functions must use keys with names that use subsequent values of N. You must specify functions in SQL syntax and should use the variables x1, x2, ..., xn to refer to the 1st, 2nd, and nth independent variables, respectively. For example,'function1' -> 'sin(x1 * x2 + x3)', 'function2' -> 'cos(x1 * x3)'.
Optional
convergenceTolerance — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive double representing the relative convergence threshold for the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) optimizer. The optimizer stops when the relative change in coefficients between iterations falls below this value. This option only applies when you set the lassoCoefficient option to a positive value. The default value is 0.000001 (1e-6).
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
fitIntercept — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option to false, the model does not fit a y-intercept, forcing the regression through the origin. This option is equivalent to setting yIntercept to 0, and you cannot combine them. The default value is true.
gamma — If you set this option, the value must be a matrix. This value represents a Tikhonov gamma matrix used for regularization. For details, see Tikhonov regularization.
huberConvergenceTolerance — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive double type that represents the convergence tolerance for Huber iteratively reweighted least squares (IRLS) iterations. The algorithm stops when the maximum coefficient change between iterations falls below this value; if the algorithm reaches huberMaxIterations first, training succeeds with a warning logged but no error returned. This option only applies when you set the lossFunction option to 'huber'. The default value is 1e-6.
huberDelta — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive double type that represents the Huber loss threshold expressed as a multiple of the residual scale (σ). Residuals smaller than delta * σ use squared loss; residuals larger use linear loss. The model estimates σ from the initial OLS fit and refreshes it each IRLS iteration via the weighted residual standard deviation, so a delta of 1.35 matches the convention used by scikit-learn’s HuberRegressor.epsilon and Spark’s HuberAggregator. This option only applies when you set the lossFunction option to 'huber'. The default value is 1.35.
huberMaxIterations — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer type that represents the maximum number of Huber IRLS iterations. If the algorithm has not converged within this many iterations, training succeeds with a warning logged but no error returned. This option only applies when you set the lossFunction option to 'huber'. The default value is 20.
lassoCoefficient — (BETA. This option is part of an in-development feature and is subject to change.) If you specify this option, the value must be a nonnegative double. This option adds L1 (lasso) regularization, which encourages sparse coefficients by driving some toward exactly zero. When you set this option to a positive value, the model uses the FISTA optimizer instead of the closed-form normal equation. When normalize=true (the default), the penalty is applied in standardized space and the persisted coefficients are un-scaled to the original units, matching scikit-learn’s Lasso(normalize=True) convention. You can combine this option with the ridgeCoefficient option for elastic net regularization. You cannot combine this option with the gamma option. The default value is 0.0 (no L1 regularization).
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
lossFunction — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be either 'squared_error' or 'huber'. When you set this value to 'huber', the model uses Huber loss with iteratively reweighted least squares (IRLS), which is robust to outliers. Huber loss cannot be combined with the lassoCoefficient, gamma, or weighted options, and forces normalize to false because IRLS requires raw-scale residuals. Huber loss IS compatible with the ridgeCoefficient and fitIntercept options. The default value is 'squared_error'.
maxIterations — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer type that represents the maximum number of FISTA iterations. This option only applies when you set the lassoCoefficient option to a positive value. The default value is 1000.
metrics — If you set this option to true, the model collects quality metrics such as the coefficient of determination (R-squared), the adjusted coefficient of determination, and the root mean squared error (RMSE). The default value is false.
normalize — If you set this option to true, the model uses auto-scaling to compute the mean and standard deviation of each input feature to normalize data during training, making training more numerically stable. The model then unscales parameters so the persisted model operates in the original units. The default value is true.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
ridgeCoefficient — (BETA. This option is part of an in-development feature and is subject to change.) If you specify this option, the value must be a nonnegative double. This option adds L2 (ridge) regularization to the normal equation by adding this value times the identity matrix to the X-transpose-X (Gram) matrix before inversion. Larger values shrink the coefficients more toward zero. When normalize=true (the default), the penalty is applied in standardized space and the persisted coefficients are un-scaled to the original units, matching scikit-learn’s Ridge(normalize=True) convention. The default value is 0.0 (no regularization). You can combine this option with the gamma option or with the lassoCoefficient option for elastic net regularization.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
threshold — This option enables soft thresholding. If you specify this option, the option must be a positive numeric value. After the model calculates the coefficients, if any coefficients are greater than the threshold value, the model subtracts the threshold value from the coefficients. If any coefficients are less than the negation of the threshold value, the model adds the threshold value to the coefficients. For any coefficients between the negative and positive threshold values, the model sets the coefficients to zero.
weighted — If you set this option to true, the model performs weighted least squares regression, where each sample has an associated weight or importance. When weighted, there is an extra numeric column after the dependent variable that represents the weight of the sample. The default value is false.
yIntercept — If you set this option, then the option must be a numeric value. The system forces the specific y-intercept (i.e., the model value when x is zero).
Linear Discriminant Analysis
Model Options
Optional
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
normalize — If you set this option to true, the model automatically computes the mean and standard deviation of each feature and uses them to normalize the data during training. The default value is true.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
Logistic Regression
Model Options
Optional
ROCNumSamples — If you set this option, you must specify a positive integer that represents the number of samples for the model to use when calculating the area under the ROC curve. You must also set the metrics option to true. The default value is 10.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
metrics — If you set this option to true, the model calculates the percentage of samples that are correctly classified by the model and saves this information in the sys.logistic_regression_models system catalog table. The default value is false.
normalize — If you set this option to true, the model uses auto-scaling to compute the mean and standard deviation of each input feature to normalize data during training, making training more numerically stable. The model then unscales parameters so the persisted model operates in the original units. The default value is true.
numEpochs — If you set this option, the value must be a positive integer type representing the maximum number of IRLS iterations during training. The default value is 20.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
ridgeCoefficient — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a non-negative double type representing the L2 regularization strength. The model adds this value times an identity matrix to the Hessian diagonal during IRLS optimization, shrinking coefficients toward zero and reducing overfitting. The bias (intercept) term is not regularized. The default value is 0.0 (no regularization).
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
weighted — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option to true, the model performs weighted logistic regression, where each sample has a weight or importance associated with it. In this case, the table includes an additional numeric column after the dependent variable that contains the sample weight. If you set this option to auto, the model automatically computes balanced class weights using the formula weight_k = max_count / count_k, where max_count is the count of the most-frequent class. The most-frequent class receives weight 1.0 and minority classes receive proportionally larger weights; relative weighting between classes is identical to class_weight='balanced' from scikit-learn (only the absolute scale differs). The model applies the weight as a multiplier on the per-sample log-loss term during IRLS. The default value is false.
Mlp
Model Options
Required
hiddenLayerSize — You must set this option to a positive integer type that specifies the number of nodes in each hidden layer.
hiddenLayers — You must set this option to a positive integer type that specifies how many hidden layers to use.
lossFunction — This option specifies the loss function that all hidden layer nodes and all output layer nodes use. This function can be one of several predefined loss functions or a user-defined loss function. The predefined loss functions are squared_error (regression), vector_squared_error (vector-valued regression), log_loss (binary classification with target values of 0 and 1), logits_loss (binary classification with target values of 0 and 1), hinge_loss (binary classification with target values of -1 and 1), and cross_entropy_loss (multi-class classification). If the value for this required option is none of these strings, the model assumes a user-defined loss function. The user-defined loss function specifies the per-sample loss. Then, the actual loss function is the sum of this function applied to all samples. The model should use the variable y to refer to the dependent variable in the training data, and the model should use the variable f to refer to the computed estimate for the specified sample.
outputs — You must set this option to a positive integer that specifies the number of outputs.
Optional
activationFunction — If you set this option, the values are linear, relu (rectified linear unit), leakyrelu (leaky rectified linear unit), tanh (hyperbolic tangent function), or sigmoid (fast sigmoid approximation). The default value is relu. This option affects all layers except the output layer.
adamBeta1 — If you set this option, the option represents the value of β₁ in the Adam optimization algorithm. For higher values of this option, training is less noisy but takes longer to converge. The default value is 0.9.
adamBeta2 — If you set this option, the option represents the value of β₂ in the Adam optimization algorithm. For higher values of this option, training is less noisy but takes longer to converge. The default value is 0.99.
adamEpsilon — If you set this option, the option represents the value of ε in the Adam optimization algorithm. For higher values of this option, training is more numerically stable but takes longer to converge. The default value is 1e-7.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
finiteDifferenceH — If you set this option, the value must be a double representing the step size (h) for approximating gradients using the finite difference method. The model uses this value only if analytical gradients are not active. This value should generally be a small positive value, typically from 0.0001 (1e-4) to 0.0000001 (1e-7). The default value is 0.00001 (1e-5).
gradientClipThreshold — If you set this option, the value must be a double that represents the gradient norm threshold for clipping. When the overall gradient norm exceeds this threshold, the system scales all gradient components uniformly to preserve direction. This operation prevents issues with exploding gradients in unstable loss landscapes. Set this value to 0 or a negative value to disable gradient clipping. The default value is 1000000 (1e6).
learningRate — If you set this option, the value must be a double type representing the base learning rate for the Adam (Adaptive Moment Estimation) machine learning optimizer. Adam adapts this rate individually for each parameter during training. A common starting point for Adam is 0.001 (1e-3). Valid values must be positive and are generally in the range of 0.00001 (1e-5) to 0.01 (1e-2). A higher learning rate can speed up training, but can cause the optimizer to overshoot and miss optimal solutions. Conversely, a lower learning rate ensures more stable and precise convergence but can make training much slower. If you do not specify this option, the system automatically selects a learning rate and adjusts it during training using the 1Cycle learning schedule. Specifying a learning rate disables automatic adjustment and instead uses a fixed learning rate value.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
maxInitParamValue — If you set this option, the value must be a floating-point number. Sets the maximum for initial parameter values in the optimization algorithm. The default value is 1.
metrics — If you set this option to true, the model calculates the average value of the loss function.
minInitParamValue — If you set this option, the value must be a floating-point number. Sets the minimum for initial parameter values in the optimization algorithm. The default value is -1.
normalize — If you set this option to true, this option applies z-score normalization to inputs by default, storing means and standard deviation, and automatically applying them at inference. The default value is true.
numEpochs — If you set this option, the value must be a positive integer type representing the maximum number of epochs, or full passes, during training through the entire data set. If you do not specify this option, the default maximum is 200, but training typically stops earlier due to automatic early stopping when the model has converged.
outputActivationFunction — If you set this option, the values are linear, relu (rectified linear unit), leakyrelu (leaky rectified linear unit), tanh (hyperbolic tangent function), or sigmoid (fast sigmoid approximation). Different activation functions have different output ranges. The chosen activation function should match the dependent variable of your data. For example, if the dependent variable can be anything, then choose the linear value. If the dependent variable is always positive, then choose the relu value. If your outputs range from -1 to 1 or you perform hinge loss classification, tanh is a good option because the hyperbolic tangent function has the same range. But, if your outputs range from 0 to 1 or you perform log loss classification, sigmoid is a better choice for the same reason. This option defaults to linear. The option only sets the activation function for the output layer.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
randomSeed — If you set this option, the value must be a positive integer type representing the seed for the random number generator the system uses for weight initialization. Setting this option makes model training deterministic (given the same data and options). If you do not specify this option or set it to 0, the system uses a non-deterministic random seed.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
useSoftmax — If you set this option to true, the model applies the softmax function to the output of the output layer before computing the loss function. The default value is true if you set the lossFunction option to cross_entropy_loss, and false otherwise.
Multiple Linear Regression
Model Options
Optional
convergenceTolerance — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive double type that represents the relative convergence threshold for the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) optimizer. The optimizer stops when the relative change in coefficients between iterations falls below this value. This option only applies when you set the lassoCoefficient option to a positive value. The default value is 0.000001 (1e-6).
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
fitIntercept — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option to false, the model does not fit a y-intercept, forcing the regression through the origin. This option is equivalent to setting yIntercept to 0, and you cannot combine them. The default value is true.
gamma — If you set this option, the option must be a matrix. The value represents a Tikhonov gamma matrix used for regularization. For details, see Tikhonov regularization. The model uses this option for ridge regression.
huberConvergenceTolerance — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive double type that represents the convergence tolerance for Huber iteratively reweighted least squares (IRLS) iterations. The algorithm stops when the maximum coefficient change between iterations falls below this value; if the algorithm reaches huberMaxIterations first, training succeeds with a warning logged but no error returned. This option only applies when you set the lossFunction option to 'huber'. The default value is 1e-6.
huberDelta — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive double type that represents the Huber loss threshold expressed as a multiple of the residual scale (σ). Residuals smaller than delta * σ use squared loss; residuals larger use linear loss. The model estimates σ from the initial OLS fit and refreshes it each IRLS iteration via the weighted residual standard deviation, so a delta of 1.35 matches the convention used by scikit-learn’s HuberRegressor.epsilon and Spark’s HuberAggregator. This option only applies when you set the lossFunction option to 'huber'. The default value is 1.35.
huberMaxIterations — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer type that represents the maximum number of Huber IRLS iterations. If the algorithm has not converged within this many iterations, training succeeds with a warning logged but no error returned. This option only applies when you set the lossFunction option to 'huber'. The default value is 20.
lassoCoefficient — (BETA. This option is part of an in-development feature and is subject to change.) If you specify this option, the value must be a nonnegative double. This option adds L1 (lasso) regularization, which encourages sparse coefficients by driving some toward exactly zero. When you set this option to a positive value, the model uses the FISTA optimizer instead of the closed-form normal equation. When normalize=true (the default), the penalty is applied in standardized space and the persisted coefficients are un-scaled to the original units, matching scikit-learn’s Lasso(normalize=True) convention. You can combine this option with the ridgeCoefficient option for elastic net regularization. You cannot combine this option with the gamma option. The default value is 0.0 (no L1 regularization).
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
lossFunction — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be either 'squared_error' or 'huber'. When you set this value to 'huber', the model uses Huber loss with iteratively reweighted least squares (IRLS), which is robust to outliers. Huber loss cannot be combined with the lassoCoefficient, gamma, or weighted options, and forces normalize to false because IRLS requires raw-scale residuals. Huber loss IS compatible with the ridgeCoefficient and fitIntercept options. The default value is 'squared_error'.
maxIterations — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer type that represents the maximum number of FISTA iterations. This option only applies when the lassoCoefficient option is set to a positive value. The default value is 1000.
metrics — If you set this option to true, the model collects quality metrics such as the coefficient of determination (R-squared), the adjusted coefficient of determination, and the root mean squared error (RMSE). The default value is false.
normalize — If you set this option to true, the model uses auto-scaling to compute the mean and standard deviation of each input feature to normalize data during training, making training more numerically stable. The model then unscales parameters so the persisted model operates in the original units. The default value is true.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
ridgeCoefficient — (BETA. This option is part of an in-development feature and is subject to change.) If you specify this option, the value must be a nonnegative double. This option adds L2 (ridge) regularization to the normal equation by adding this value times the identity matrix to the X-transpose-X (Gram) matrix before inversion. Larger values shrink the coefficients more toward zero. When normalize=true (the default), the penalty is applied in standardized space and the persisted coefficients are un-scaled to the original units, matching scikit-learn’s Ridge(normalize=True) convention. The default value is 0.0 (no regularization). You can combine this option with the gamma option or with the lassoCoefficient option for elastic net regularization.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
threshold — If you set this option, the option enables soft thresholding. The value must be a positive number. After the model calculates the coefficients, if any coefficients exceed the threshold value, the model subtracts the threshold value from those coefficients. If any coefficients are less than the negation of the threshold value, the model adds the threshold value to the coefficients. For any coefficients that are between the negative and positive threshold values, the model sets the coefficients to zero.
weighted — If you set this option to true, the model performs weighted least squares regression, where each sample has a weight or importance associated with it. In this case, the table contains an additional numeric column after the dependent variable, which contains the weight for the sample. The default value is false.
yIntercept — If you set this option, then the option must be a numeric value. The system forces the specific y-intercept (i.e., the model value when x is zero).
Naive Bayes
Model Options
Optional
continuousFeatures — If you set this option, the value must be a comma-separated list of the feature indexes that are continuous numeric variables. Indexes start with 1. In the default state, the model considers no features as continuous.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
metrics — If you set this option to true, the model calculates the percentage of samples correctly classified by the model and saves this information in a system catalog table. The default value is false.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
Nonlinear Regression
Model Options
Required
function — Specify the name of the function to fit the data in SQL syntax. Use a1, a2, … to refer to the parameters for optimization. Use x1, x2, … to refer to the input features. The model does not allow some SQL functions. The model allows only scalar expressions that can be represented internally as postfix expressions. Most notably, the model does not allow some functions that are rewritten as CASE statements (like least() and greatest()). If your function is not allowed, the model displays an error message.
numParameters — Specify this option as a positive integer. This value specifies the number of different parameters to optimize, i.e., how many different aN variables there are in the user-specified function.
Optional
adamBeta1 — If you set this option, the option represents the value of β₁ in the Adam optimization algorithm. For higher values of this option, training is less noisy but takes longer to converge. The default value is 0.9.
adamBeta2 — If you set this option, the option represents the value of β₂ in the Adam optimization algorithm. For higher values of this option, training is less noisy but takes longer to converge. The default value is 0.99.
adamEpsilon — If you set this option, the option represents the value of ε in the Adam optimization algorithm. For higher values of this option, training is more numerically stable but takes longer to converge. The default value is 1e-7.
constraints — If you set this option, the value must be a SQL Boolean predicate over the parameter variables a1, a2, …, aN. During training, the model rejects every proposed parameter step that would violate the predicate. The LM optimizer increases its damping factor, and the Adam optimizer skips the update. The predicate can use any Boolean SQL composition (AND, OR, NOT, BETWEEN, comparisons) but cannot reference feature variables (x1, x2, …) or the dependent variable (y). Referencing these variables causes the CREATE MLMODEL SQL statement to fail with a validation error. If the optimizer rejects the number of proposed steps that you specify using the maxConsecutiveRejections option in a row without an accept, training terminates with constraint_infeasible = true in sys.nonlinear_regression_models instead of converging. The system persists the constraint string with the model and shows it in the sys.nonlinear_regression_models system catalog table.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
finiteDifferenceH — If you set this option, the value must be a double type representing the step size (h) for approximating gradients using the finite difference method. The model uses this value only if analytical gradients are not active. This value should generally be a small positive number, typically from 0.0001 (1e-4) to 0.0000001 (1e-7). The default value is 0.00001 (1e-5).
gradientClipThreshold — If you set this option, the value must be a double that represents the gradient norm threshold for clipping. When the overall gradient norm exceeds this threshold, the system scales all gradient components uniformly to preserve direction. This operation prevents issues with exploding gradients in unstable loss landscapes. Set this value to 0 or a negative value to disable gradient clipping. The default value is 1000000 (1e6).
lassoCoefficient — If you specify this option, the value must be a double data type. This option is the lasso coefficient for the loss function. The default behavior is the function ignores this option, effectively setting this option to 0.0.
learningRate — If you set this option, the value must be a double type representing the base learning rate for the Adam (Adaptive Moment Estimation) machine learning optimizer. Adam adapts this rate individually for each parameter during training. A common starting point for Adam is 0.001 (1e-3). Valid values must be positive and are generally in the range of 0.00001 (1e-5) to 0.01 (1e-2). A higher learning rate can speed up training, but can cause the optimizer to overshoot and miss optimal solutions. Conversely, a lower learning rate ensures more stable and precise convergence but can make training much slower. If you do not specify this option, the system automatically selects a learning rate and adjusts it during training using the 1Cycle learning schedule. Specifying a learning rate disables automatic adjustment and instead uses a fixed learning rate value.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
logRejections — If you set this option to true, the database emits an INFO-level log line for every constraint-rejected step during training. This option only takes effect when you set the constraints option. The default value is false.
lossFunction — If you set this option, the option indicates to the nonlinear optimizer the loss function to use on a per-sample basis. Then, the actual loss function is the sum of this function applied to all samples. The model should use the variable y to refer to the dependent variable in the training data and the variable f to refer to the computed estimate for the specified sample. The default is the least squares function, which you can specify as (f-y)*(f-y).
maxConsecutiveRejections — If you set this option, the value must be a positive integer type representing the number of consecutive constraint-rejected steps that trigger termination during training with constraint_infeasible = true in sys.nonlinear_regression_models. This option only takes effect when you set the constraints option. Larger values let the optimizer probe further along a tightly-binding constraint boundary; smaller values surface infeasible problems faster. The accept and reject counter resets to zero on every accepted step. The default value is 50.
maxInitParamValue — If you specify this option, the value must be a floating-point number. This option sets the maximum for initial parameter values in the optimization algorithm. The default value is 1.
metrics — If you set this option to true, the model calculates the coefficient of determination (R-squared), the adjusted R-squared, and the root mean squared error (RMSE). However, the model calculates these quality metrics using the least squares loss function, and not the user-specified loss function, because these metrics only make sense for least squares. The default value is false.
minInitParamValue — If you specify this option, the value must be a floating-point number. This option sets the minimum for initial parameter values in the optimization algorithm. The default value is -1.
numEpochs — If you set this option, the value must be a positive integer type representing the maximum number of epochs, or full passes, during training through the entire data set. If you do not specify this option, the default maximum value is 200 for Adam optimization or 100 for Levenberg-Marquardt, but training typically stops earlier due to automatic early stopping when the model has converged.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
randomSeed — If you set this option, the value must be a positive integer type representing the seed for the random number generator the system uses for weight initialization. Setting this option makes model training deterministic (given the same data and options). If you do not specify this option or set it to 0, the system uses a non-deterministic random seed.
ridgeCoefficient — If you specify this option, the value must be a double data type. This option is the ridge coefficient for the loss function. The default behavior is the function ignores this option, effectively setting this option to 0.0.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
Polynomial Regression
Model Options
Required
order — This option is the degree of the polynomial and must be set to a positive integer.
Optional
convergenceTolerance — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive double type that represents the relative convergence threshold for the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) optimizer. The optimizer stops when the relative change in coefficients between iterations falls below this value. This option only applies when you set the lassoCoefficient option to a positive value. The default value is 0.000001 (1e-6).
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
fitIntercept — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option to false, the model does not fit a y-intercept, forcing the regression through the origin. This option is equivalent to setting yIntercept to 0, and you cannot combine them. The default value is true.
gamma — If you specify this option, the value must be a matrix. The value represents a Tikhonov gamma matrix that is used for regularization. For details, see Tikhonov regularization.
huberConvergenceTolerance — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive double type that represents the convergence tolerance for Huber iteratively reweighted least squares (IRLS) iterations. The algorithm stops when the maximum coefficient change between iterations falls below this value; if the algorithm reaches huberMaxIterations first, training succeeds with a warning logged but no error returned. This option only applies when you set the lossFunction option to 'huber'. The default value is 1e-6.
huberDelta — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive double type that represents the Huber loss threshold expressed as a multiple of the residual scale (σ). Residuals smaller than delta * σ use squared loss; residuals larger use linear loss. The model estimates σ from the initial OLS fit and refreshes it each IRLS iteration via the weighted residual standard deviation, so a delta of 1.35 matches the convention used by scikit-learn’s HuberRegressor.epsilon and Spark’s HuberAggregator. This option only applies when you set the lossFunction option to 'huber'. The default value is 1.35.
huberMaxIterations — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer type that represents the maximum number of Huber IRLS iterations. If the algorithm has not converged within this many iterations, training succeeds with a warning logged but no error returned. This option only applies when you set the lossFunction option to 'huber'. The default value is 20.
interactionOnly — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option to true, the model includes only interaction terms in the polynomial expansion. Interaction terms are products of distinct features where no single feature has a power greater than 1. For example, with three features and order 2, the model includes x1*x2 and excludes x1^2. The default value is false.
lassoCoefficient — (BETA. This option is part of an in-development feature and is subject to change.) If you specify this option, the value must be a nonnegative double. This option adds L1 (lasso) regularization, which encourages sparse coefficients by driving some toward exactly zero. When you set this option to a positive value, the model uses the FISTA optimizer instead of the closed-form normal equation. When normalize=true (the default), the penalty is applied in standardized space and the persisted coefficients are un-scaled to the original units, matching scikit-learn’s Lasso(normalize=True) convention. You can combine this option with the ridgeCoefficient option for elastic net regularization. You cannot combine this option with the gamma option. The default value is 0.0 (no L1 regularization).
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
lossFunction — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be either 'squared_error' or 'huber'. When you set this value to 'huber', the model uses Huber loss with iteratively reweighted least squares (IRLS), which is robust to outliers. Huber loss cannot be combined with the lassoCoefficient, gamma, or weighted options, and forces normalize to false because IRLS requires raw-scale residuals. Huber loss IS compatible with the ridgeCoefficient and fitIntercept options. The default value is 'squared_error'.
maxIterations — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer type that represents the maximum number of FISTA iterations. This option only applies when you set the lassoCoefficient option to a positive value. The default value is 1000.
metrics — If you set this option to true, the model collects quality metrics such as the coefficient of determination (R-squared), the adjusted coefficient of determination, and the root mean squared error (RMSE). The default value is false.
negativePowers — If you set this option to true, the model includes independent variables raised to negative powers. These variables are named Laurent polynomials. The model generates all possible terms such that the sum of the absolute value of the power of each term in each product is less than or equal to the order. For example, with two independent variables and the order set to 2, the model is: y = a1*x1^2 + a2*x1^-2 + a3*x2^2 + a4*x2^-2 + a5*x1*x2 + a6*x1^-1*x2 + a7*x1*x2^-1 + a8*x1^-1*x2^-1 + a9*x1 + a10*x1^-1 + a11*x2 + a12*x2^-1 + b. The default value is false.
normalize — If you set this option to true, the model uses auto-scaling to compute the mean and standard deviation of each input feature to normalize data during training, making training more numerically stable. The model then unscales parameters so the persisted model operates in the original units. The default value is true.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
ridgeCoefficient — (BETA. This option is part of an in-development feature and is subject to change.) If you specify this option, the value must be a nonnegative double. This option adds L2 (ridge) regularization to the normal equation by adding this value times the identity matrix to the X-transpose-X (Gram) matrix before inversion. Larger values shrink the coefficients more toward zero. When normalize=true (the default), the penalty is applied in standardized space and the persisted coefficients are un-scaled to the original units, matching scikit-learn’s Ridge(normalize=True) convention. The default value is 0.0 (no regularization). You can combine this option with the gamma option or with the lassoCoefficient option for elastic net regularization.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
threshold — This option enables soft thresholding. If you specify this option, then the option must be a positive numeric value. After the model calculates the coefficients, if any of them are greater than the threshold value, the threshold value is subtracted from them. If any coefficients are less than the negation of the threshold value, the model adds the threshold value to them. For any coefficients that are between the negative and positive threshold values, the model sets those coefficients to zero.
weighted — If you set this option to true, the model performs weighted least squares regression, where each sample has an associated weight. When weighted, there is an extra numeric column after the dependent variable that has the weight for the sample. The default value is false.
yIntercept — If you set this option, then the option must be a numeric value. The system forces the specific y-intercept (i.e., the model value when x is zero).
Principal Component Analysis
Model Options
Optional
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
numComponents — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer specifying the number of principal components to retain. This value must be between 1 and the number of input features. If you do not set this option or set it to 0, the model retains all components. This model retains components in order of decreasing explained variance.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
whiten — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option to true, the model scales each component score by the inverse square root of its eigenvalue, producing unit-variance outputs. This option is useful when downstream algorithms are sensitive to feature scale. The default value is false.
Random Forest
Model Options
Required
numChildren — Number of child decision trees.
Optional
ROCNumSamples — If you set the option, you must also set the metrics option. This positive integer indicates the number of samples for the model to use for the area under the ROC curve. The default value is the number of child decision trees.
bootstrap — If you set this option to true, the model uses bootstrap sampling with replacement, meaning the model trains each tree in the random forest on a random subset of the data (either the rowsPerChild or fractionSelected option sets the exact number of rows), and the same row can appear multiple times in each tree. If you set this option to false, this option does not use replacement, meaning each row can appear at most once per tree. The default value is false.
continuousFeatures — If you set this option, the value must be a comma-separated list of the feature indexes that are continuous numeric variables. Indexes start with 1. In the default state, the model considers no features as continuous.
distinctCountLimit — If you set this option, the value must be a positive integer. This value limits how many distinct values a non-continuous feature and the label can contain. The default value is 256.
doPrune — If you set this option to true, the model uses Pessimistic Error Pruning (PEP) to prune the tree after training. The default value is false.
enableResplits — If you set this option, the value must be a boolean type that determines if the tree can reuse the same continuous feature multiple times along a single branch (e.g., split on x1 < 7 and later x1 < 3). This action can capture more complex, range-specific relationships. The default value is true, meaning that continuous features remain available for additional splits after use, thereby allowing the tree to create more complex decision boundaries. If you set this option to false, the model marks continuous features as exhausted after their first use, and the model cannot use them again in subsequent splits in the same tree. The model passes this option directly to the child decision trees.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
featureSubsetStrategy — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the model passes this option directly to the child decision trees. The option specifies how many features the child decision trees should consider at each split from the still-available features. When this value is higher, the model has higher accuracy and lower variance, but takes longer to train. You can specify this option either as an integer (e.g., 4, meaning consider up to four features at each split) or one of these values: all (checks every feature), sqrt (checks up to the square root of the number of total features), log2 (checks up to the base-2 logarithm of the number of total features), and one-third (checks up to one-third of the number of total features). The default value is all.
fractionSelected — If you set this option, the option represents the proportion of rows the model uses to train each child model. The value is a double that must be in the interval (0, 1]. You cannot set this option if you also set the rowsPerChild option to a positive value. The default behavior is that the model uses all available rows.
inputsPerChild — If you set this option, the option specifies the number of features for the creation of each child decision tree. The default value is the number of features you specify for the forest divided by 3 and rounded up.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
maxCellsToFetch — If you set this option, the model passes this option directly to the child decision trees. If you set this option, the value must be a positive integer. Controls the chunking behavior when fetching feature values during model training. The limit represents the maximum number of data cells (calculated as the number of columns × number of rows) that the system can fetch in a single operation, not a byte limit. When the expected data size exceeds this threshold, the algorithm switches to database-based processing using SQL queries instead of in-memory processing. The default value is 33,554,432 cells (calculated as 32 × 1024 × 1024).
maxChildThreads — If you set this option, the value must be an integer type representing the maximum number of threads each child decision tree can use. The default value is 1.
maxDepth — If you set this option, the value must be a positive integer. This value sets the maximum allowable depth of the decision tree. The default value is 3.
maxLeafNodes — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be an integer greater than or equal to 2. The model passes this option directly to the child decision trees. The value caps the number of leaf nodes in each child tree during its depth-first build (it is a soft cap rather than best-first growth). The default value is unspecified, which means there is no limit on the number of leaf nodes.
maxThreads — If you set this option, the option specifies the maximum number of parallel threads to use while the model trains decision trees. This value must be a positive integer. The default value is 16.
metrics — If you set this option to true, the model also calculates the percentage of samples that are correctly classified by the model for the random forest and saves this information in a system catalog table. The default value is false.
minImpurityDecrease — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a non-negative double. The child decision tree splits only if the decrease in impurity is greater than or equal to this value. The model passes this option directly to the child decision trees. The default value is 0.0.
minSamplesLeaf — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer. This value sets the minimum number of samples required in a leaf node. The model passes this option directly to the child decision trees. The default value is 1.
minSamplesSplit — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be an integer greater than or equal to 2. This value sets the minimum number of samples required to split a node. The model passes this option directly to the child decision trees. The default value is 2.
numSplits — If you set this option, the value must be an integer type greater than 1. This value sets the maximum number of binary branches a continuous feature can consider. The default value is 32.
oobScore — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option to true, the model computes the Out-of-Bag (OOB) score after training. The OOB score estimates model accuracy using only the samples not included in the bootstrap sample of each tree. This score provides a built-in validation metric without requiring a separate test set. This option requires that you set the bootstrap option to true as well. The default value is false.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
requiredFeatures — If you set this option, the option must be a comma-separated list of integers as strings representing specific features where the first feature has the value 1. The model uses these features in every decision tree in the forest. The default behavior is that the decision tree in the forest can train on any feature in the list.
rowsPerChild — If you set this option to a positive integer, the number represents the number of rows (from a random sample) to use for each decision tree. If you set this option to 0, each child uses all available rows. The default value is 0. You cannot set this option to a positive value if you also set the fractionSelected option.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
skipLimitCheck — If you set this option to true, the model skips cardinality checks that throw errors when columns have too many values. The limit that this option checks is the same one specified by the distinctCountLimit option. This option defaults to false.
splitMetric — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the option controls which function the model uses to evaluate the quality of a split during tree construction. The model passes this option directly to the child decision trees. Supported options are: gini_impurity (measures impurity based on class distributions), entropy (measures impurity using information entropy), and hellinger (uses the Hellinger distance, which requires exactly two output classes). The default value is gini_impurity.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
weighted — If you set this option, the model considers weights for labels. If you set this option value to true, you must specify an additional column as a double in the training data for label weights. Rows with the same labels must have the same weights. If you set this value to auto, the model calculates weights automatically by weighting each label according to the ratio of the count of the most frequent label to the count of the specified label. As a result, the most frequent label has a weight of 1.0, and the other label weights are higher. This option defaults to false, which means all labels have equal weight.
Regression Tree
Model Options
Optional
continuousFeatures — If you set this option, the value must be a comma-separated list of the feature indexes that are continuous numeric variables. Indexes start with 1. In the default state, the model considers no features as continuous.
distinctCountLimit — If you set this option, the value must be a positive integer. This value sets the limit for the number of distinct values a non-continuous feature and the label can contain. This option defaults to 256.
enableResplits — If you set this option, the value must be a boolean type that determines if the tree can reuse the same continuous feature multiple times along a single branch (e.g., split on x1 < 7 and later x1 < 3). This action can capture more complex, range-specific relationships. The default value is true, meaning that continuous features remain available for additional splits after use, which allows the tree to create more complex decision boundaries. When you set this option to false, the model marks continuous features as exhausted after their first use, and the model cannot use them again in subsequent splits in the same tree.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
featureSubsetStrategy — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the option specifies how many features the regression tree should consider at each split from the still-available features. When this value is higher, the model has a higher accuracy and lower variance, but takes longer to train. You can specify this option either as an integer (e.g., 4, meaning consider up to four features at each split) or one of these values: all (check every feature), sqrt (check up to the square root of the number of total features), log2 (check up to the base-2 logarithm of the number of total features), and one-third (check up to one-third of the number of total features). The default value is all.
maxCellsToFetch — If you set this option, the value must be a positive integer. Controls the chunking behavior when fetching feature values during model training. The limit represents the maximum number of data cells (calculated as the number of columns × number of rows) that the system can fetch in a single operation, not a byte limit. When the expected data size exceeds this threshold, the algorithm switches to database-based processing using SQL queries instead of in-memory processing. The default value is 33,554,432 cells (calculated as 32 × 1024 × 1024).
maxDepth — If you set this option, the value must be a positive integer. This value sets the maximum allowable depth of the decision tree (the maximum number of features to split on). The default is unspecified, which means there is no maximum depth.
maxLeafNodes — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be an integer greater than or equal to 2. This value caps the number of leaf nodes in the regression tree during the standard depth-first build. When the tree reaches this limit, the model stops queuing further splits and converts pending nodes to leaves; this is a soft cap rather than the best-first growth, so the resulting tree might differ from a best-first tree of the same leaf budget. The default value is unspecified, which means there is no limit on the number of leaf nodes.
maxRows — If you set this option, the value must be a positive integer. This option limits the number of rows used for model training by creating a snapshot table with only the specified number of rows from the input query. When this option is unspecified, the model trains using all rows from the input query.
maxThreads — If you set this option, the value must be a positive integer. This value indicates the maximum number of parallel threads to use while the model trains. The default value is 2.
metrics — If you set this option to true, the model also calculates the percentage of samples correctly classified by the model and saves this information in a system catalog table. The default value is false.
minImpurityDecrease — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a non-negative double. The tree splits only if the decrease in impurity (the difference between the parent node impurity and the weighted sum of child node impurities, measured by variance) is greater than or equal to this value. The default value is 0.0, meaning the model accepts all splits that reduce variance.
minSamplesLeaf — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer. This value sets the minimum number of samples required in a leaf node. The model rejects any split that creates a child node with fewer samples than this value, making the node a leaf instead. The default value is 1, meaning every leaf must have at least one sample.
minSamplesSplit — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be an integer greater than or equal to 2. This value sets the minimum number of samples required to split a node. Nodes with fewer samples than this value become leaves. The default value is 2.
numSplits — If you set this option, the value must be an integer greater than 1. This value sets the maximum number of binary branches a continuous feature can consider. The default value is 32.
queryInternalParallelism — If you set this option, the database appends the USING PARALLELISM = <value> clause to all intermediate SQL queries the model executes during training, where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
resplitDepth — If you set this option, the value must be an integer type that sets the maximum depth at which tree nodes can be re-split during optimization. Controls how deep the algorithm searches for better split points. The default value is 6.
resplitThreshold — If you set this option, the value must be a decimal type that sets the minimum improvement threshold required to trigger a re-split operation. Lower values (e.g., 0.01) allow more aggressive re-splitting but can increase training time. Higher values (e.g., 1.0) require larger improvements to trigger re-splits. The default value is 0.1.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
skipLimitCheck — If you set this option to true, the model skips cardinality checks that throw errors when columns have too many values. The limit that this option checks is the same one that you specify using the distinctCountLimit option. The default value is false.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
Simple Linear Regression
Model Options
Optional
fitIntercept — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option to false, the model does not fit a y-intercept, forcing the regression through the origin. This option is equivalent to setting yIntercept to 0, and you cannot combine them. The default value is true.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
metrics — If you set this option to true, the model collects quality metrics such as the coefficient of determination (R-squared) and the root mean squared error (RMSE). The default value is false.
normalize — If you set this option to true, the model uses auto-scaling to compute the mean and standard deviation of each input feature to normalize data during training, making training more numerically stable. The model then unscales parameters so the persisted model operates in the original units. The default value is true.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
ridgeCoefficient — (BETA. This option is part of an in-development feature and is subject to change.) If you specify this option, the value must be a nonnegative double. This option adds L2 (ridge) regularization to the normal equation by adding this value times the identity matrix to the X-transpose-X (Gram) matrix before inversion. Larger values shrink the coefficients more toward zero. When normalize=true (the default), the penalty is applied in standardized space and the persisted coefficients are un-scaled to the original units, matching scikit-learn’s Ridge(normalize=True) convention. The default value is 0.0 (no regularization).
threshold — This option enables soft thresholding. If you specify this option, then the option must be a positive numeric value. After the model calculates the coefficients, if any are greater than the threshold value, the threshold value is subtracted from them. If any coefficients are less than the negation of the threshold value, the model adds the threshold value to them. For any coefficients that are between the negative and positive threshold values, the model sets those coefficients to zero.
yIntercept — If you set this option, then the option must be a numeric value. The system forces the specific y-intercept (i.e., the model value when x is zero).
Stacking
Model Options
Required
levelZeroModels — This option specifies the level-0 child models of the stacking model. You must specify this value as a JSON array, where each object in the array has the four fields type (required), name, options, ignoreColumn, and extraCallArguments.
Optional
classLabels — (BETA. This option is part of an in-development feature and is subject to change.) Comma-separated list of distinct class labels for weighted stacking classification. If you set this option with the weights option, the model uses weighted voting for classification rather than weighted averaging. The model automatically detects class labels from training data when you specify the weights option, and you set the hasLabelColumn option to true.
extraColumnCount — If you set this option, the value must be an integer type that specifies how many non-feature columns there are in the input data. The default value is 0.
extraColumnsForLevelOne — If you set this option, this option specifies the extra columns to pass as input columns to the level-1 model, in addition to the level-0 outputs. This value should be a comma-separated list of integers starting at 1, where each integer refers to an extra column index (must be between 1 and the extraColumnCount value). The default behavior is to pass none of the extra columns.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
hasLabelColumn — If you set this option, the value must be a boolean type that specifies whether the input data includes a label column. The default value is true.
levelOneModel — (BETA. This option is part of an in-development feature and is subject to change.) This option specifies the level-1 child model of the stacking model. You must specify this value as a JSON object with the fields type (required), name, options, ignoreColumn, and extraCallArguments. This option is required unless you specify the weights option.
maxThreads — If you set this option, the option specifies the maximum number of parallel threads to use while the model trains. This value must be a positive integer. The default value is 16.
preservedColumnsForLevelOne — If you set this option, this option specifies the columns from the original training data to pass as an input column to the level-1 model, in addition to the level-0 outputs. This value should be a comma-separated list of integers starting at 1. The default behavior is to preserve none of the columns.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
transitionEndTime — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, you must also set the weights and transitionStartTime options. The value must be a timestamp string in the format YYYY-MM-DD HH:MM:SS. Between the values of the transitionStartTime and transitionEndTime options, the weights shift linearly from their initial values to their reversed values (e.g., [1.0, 0.0] transitions to [0.0, 1.0]). After the end time, the model uses the reversed weights permanently.
transitionStartTime — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, you must also set the weights and transitionEndTime options. The value must be a timestamp string in the format YYYY-MM-DD HH:MM:SS. Before the start time, the model uses the initial weights. Between the values of the transitionStartTime and transitionEndTime options, the weights shift linearly from their initial values to their reversed values.
weights — (BETA. This option is part of an in-development feature and is subject to change.) Comma-separated list of weights for weighted stacking. If you set this option, the model uses a weighted combination of level-zero child predictions instead of a level-one meta-learner. Weights must be non-negative and sum to 1.0, and the number of weights must match the number of level-zero children. For classification tasks, set the classLabels option or ensure that you set the hasLabelColumn option to true for automatic detection.
Support Vector Machine
Model Options
Optional
ROCNumSamples — If you set this option, you must specify a positive integer that represents the number of samples for the model to use when calculating the area under the ROC curve. You must also set the metrics option to true. The default value is 10.
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
functionN — By default, SVM uses a linear kernel. If you use a different kernel, you must provide a list of functions that are summed together, just like with linear combination regression. You must specify the first function using a key named ‘function1’. Subsequent functions must use keys with names that use subsequent values of N. You must specify functions in SQL syntax and use the variables x1, x2, … , xn to refer to the 1st, 2nd, and nth independent variables, respectively. You can specify the default linear kernel as: ‘function1’ → ‘x1’, ‘function2’ → ‘x2’, and so on. The model always adds a constant term equivalent to ‘functionN’ → ‘1.0’ that you do not need to specify explicitly.
kernel — (BETA. This option is part of an in-development feature and is subject to change.) Named kernel convenience option that generates the appropriate functionN basis expressions automatically. Supported values are: linear (default; basis is the original feature set x1, x2, ..., xn), poly (basis is every monomial of total degree 1 through kernelDegree, e.g., for two features and kernelDegree=2, the basis is x1, x2, x1*x1, x1*x2, x2*x2), rbf (basis is numRbfComponents Random Fourier Feature approximations of the form sqrt(2/D) * cos(w_j^T x + b_j) where w_j ~ N(0, 2 * kernelGamma * I) and b_j ~ Uniform(0, 2*pi)), and sigmoid (basis is per-feature tanh(kernelGamma * xi + kernelCoef0)). Note that poly, rbf, and sigmoid are feature transformations rather than inner-product kernels in the kernel-trick sense, so they are not numerically equivalent to SVC(kernel=...) from scikit-learn with the same option values. SVM cannot combine the kernel option with explicit functionN options.
kernelCoef0 — (BETA. This option is part of an in-development feature and is subject to change.) Independent term added inside the tanh of the per-feature transform of the sigmoid kernel as tanh(kernelGamma * xi + kernelCoef0). Use this option only when you set the kernel option to sigmoid. The default value is 0.0.
kernelDegree — (BETA. This option is part of an in-development feature and is subject to change.) Maximum monomial degree for the poly kernel. The generated basis includes every monomial of total degree 1 through kernelDegree (e.g., kernelDegree=2 over features x1, x2 produces x1, x2, x1*x1, x1*x2, x2*x2). Use this option only when you set the kernel option to poly. The value must be a positive integer. The default value is 3. The total number of generated basis functions grows as C(n+d, d) - 1 where n is the number of features and d is kernelDegree; the maxKernelFunctions guard rejects expansions that would exceed its limit.
kernelGamma — (BETA. This option is part of an in-development feature and is subject to change.) Coefficient for the rbf and sigmoid kernels. Use this option only when you set the kernel option to rbf or sigmoid. For rbf, controls the variance of the random projection weights w_j ~ N(0, 2 * kernelGamma * I). For sigmoid, the per-feature multiplier inside tanh(kernelGamma * xi + kernelCoef0). If you set this option to 0.0, the SVM automatically substitutes 1.0 / n_features (matching the legacy gamma='auto' behavior from scikit-learn, but not the same as the modern gamma='scale' default). The value must be a non-negative number. The default value is 0.0.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
maxKernelFunctions — (BETA. This option is part of an in-development feature and is subject to change.) Maximum number of basis functions that the kernel convenience options (poly and rbf) are allowed to generate. If the polynomial or RBF expansion would produce more functions than this limit, SVM rejects the request. The value must be a positive integer. The default value is 5000.
metrics — If you set this option to true, the model also calculates the percentage of samples that are correctly classified by the model and saves this information in a catalog table. This option defaults to false.
normalize — If you set this option to true, the model automatically computes the mean and standard deviation of each feature and uses them to normalize the data during training. Defaults to true.
numEpochs — If you set this option, the value must be a positive integer type representing the maximum number of IRLS iterations during training. The default value is 20.
numRbfComponents — (BETA. This option is part of an in-development feature and is subject to change.) Number of Random Fourier Feature components for the RBF kernel. Use this option when you set the kernel option to rbf. The value must be a positive integer. Larger values yield a more accurate approximation at the expense of training time. The default value is 100.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
randomSeed — (BETA. This option is part of an in-development feature and is subject to change.) Random seed for reproducible kernel function generation (used by the RBF kernel for Random Fourier Features). The default value is 42.
regularizationCoefficient — If you set this option, the value must be a valid floating-point number. Use this option to control the balance of finding a wide margin and minimizing incorrectly classified points in the loss function. A larger (and positive) value makes having a wide margin around the hypersurface more important relative to the incorrectly classified points. Because of how the system implements SVM, the values for this option are likely different than values used in other common SVM implementations. The default value is 1.0 / 1000000.0.
skipDropTable — If you set this option to false, the database deletes any intermediate tables created during model training. If you set this option to true, the database prevents the deletion of any intermediate tables created during model training. The default value is false.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
weighted — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option to true, the model performs weighted SVM classification, where each sample has a weight or importance associated with it. In this case, the table includes an additional numeric column after the dependent variable that contains the sample weight. If you set this option to auto, the model automatically computes balanced class weights using the formula weight_k = max_count / count_k, where max_count is the count of the most-frequent class. The most-frequent class receives weight 1.0 and minority classes receive proportionally larger weights; relative weighting between classes is identical to the class_weight='balanced' from scikit-learn (only the absolute scale differs). The model applies the weight as a multiplier on the per-sample squared-hinge loss, equivalent to per-class regularization. The default value is false.
Survival Regression
Model Options
Required
eventColumn — Specifies the name of the binary event-indicator column in the SELECT SQL statement for the training data. The column must contain 1 for an observed event and 0 for a censored observation. The column must be present in the SELECT statement with this exact name.
timeColumn — Specifies the name of the time-to-event or censoring time column in the SELECT SQL statement for the training data. The column type must be DOUBLE PRECISION, BIGINT, or TIMESTAMP. Values must be non-negative. The column must be present in the SELECT statement with this exact name.
Optional
aftDistribution — If you set this option and you set the family option to 'aft', the value selects the parametric distribution. Supported values are 'weibull' (default), 'lognormal', and 'loglogistic'. When you set the family option to 'cox', the model ignores this option. The default value is 'weibull'.
convergenceTolerance — If you set this option, the value must be a positive double type that represents the convergence threshold for the Newton-Raphson (Cox) or quasi-Newton (AFT) optimizer. The optimizer stops when the maximum absolute change in any coefficient between iterations falls below this value. The default value is 0.000001 (1e-6).
family — If you set this option, the value selects the survival regression family. Supported values are 'cox' (Cox proportional hazards) and 'aft' (accelerated failure time). The default value is 'cox'.
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
maxIterations — If you set this option, the value must be a positive integer type that represents the maximum number of Newton-Raphson (Cox) or quasi-Newton (AFT) iterations. The default value is 100.
metrics — If you set this option to true, the model computes survival metrics on the test data when you specify the VALIDATE ON clause. Otherwise, the model uses the training data. Metrics are the Concordance Index and the partial log-likelihood. The default value is false.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
regParam — If you set this option, the value must be a non-negative double type representing the L2 (ridge) regularization strength on the linear-predictor coefficients. Larger values shrink coefficients toward zero. The default value is 0.0 (no regularization).
ties — If you set this option and you set the family option to 'cox', the value selects how the model handles tied event times in the partial likelihood. Supported values are 'efron' and 'breslow'. When you set the family option to 'aft', the model ignores this option. The default value is 'efron'.
Undersampling
Model Options
Required
classColumn — (BETA. This option is part of an in-development feature and is subject to change.) The name of the column where the value defines the class for sampling. This column is required for both random and stratified strategies.
Optional
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
samplingRatio — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value indicates the final majority or minority row count ratio relative to the target class. The value must be a positive double. The default value is 1.0 (balanced output). The value of 2.0 keeps roughly twice as many non-target rows per class as the target class. Under the random strategy, sampling is non-deterministic and result counts vary by approximately the square root of the expected count. Under the stratified strategy, the per-stratum row count is deterministic, with the specific rows chosen by random ordering within each stratum.
strategy — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value indicates the sampling strategy as either random (independent Bernoulli sampling per row in each non-target class) or stratified (for each (non-target class, stratum) pair, the model keeps an exact count of rows proportional to the stratum size, with the specific rows chosen by random ordering within the stratum; this preserves the per-stratum class distribution and produces a deterministic per-stratum and overall row count). The default value is random. When you set this option to stratified, you must also set the stratumColumn option.
stratumColumn — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value sets the name of the column that defines strata. This option is required when you set the strategy option to stratified. The model preserves the per-stratum class distribution within each non-target class. The default value is unspecified.
targetClass — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value indicates the class label that defines the target row count. The default value is auto, in which case the model uses the smallest class as the target. You can also specify a class-label string, and the resulting target row count is the row count of that class.
Vector Autoregression
Model Options
Required
numLags — Specify this option as a positive integer for the number of lags in the model.
numVariables — Specify this option as a positive integer for the number of variables in the model.
Optional
convergenceTolerance — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive double type that represents the relative convergence threshold for the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) optimizer. The optimizer stops when the relative change in coefficients between iterations falls below this value. This option only applies when you set the lassoCoefficient option to a positive value. The default value is 0.000001 (1e-6).
featureArray — If you set this option to true, the model expects only one array-type input column instead of multiple columns of training data. Each array row in the input column must be the same size. The default value is false.
featureArrayElements — If you set this option, the featureArray option must be set to true. The value must be a comma-separated list of integers representing indexes of the input array to use starting at index 1. The system uses all indexes of the input array by default.
lassoCoefficient — (BETA. This option is part of an in-development feature and is subject to change.) If you specify this option, the value must be a nonnegative double. This option adds L1 (lasso) regularization, which encourages sparse coefficients by driving some toward exactly zero. When you set this option to a positive value, the model uses the FISTA optimizer instead of the closed-form normal equation. When normalize=true (the default), the penalty is applied in standardized space and the persisted coefficients are un-scaled to the original units, matching scikit-learn’s Lasso(normalize=True) convention. You can combine this option with the ridgeCoefficient option for elastic net regularization. The default value is 0.0 (no L1 regularization).
loadBalance — If you set this option, the database appends the USING load_balance_shuffle = <value> clause to all intermediate SQL queries the model executes during training where value is the specified option value (true or false). The default value is unspecified. In this case, the database does not add this clause.
maxIterations — (BETA. This option is part of an in-development feature and is subject to change.) If you set this option, the value must be a positive integer type that represents the maximum number of FISTA iterations. This option only applies when you set the lassoCoefficient option to a positive value. The default value is 1000.
metrics — If you set this option to true, the function collects the metric for the coefficient of determination (R-squared). The default value is false.
normalize — If you set this option to true, the model uses auto-scaling to compute the mean and standard deviation of each input feature to normalize data during training, making training more numerically stable. The model then unscales parameters so the persisted model operates in the original units. The default value is true.
queryInternalParallelism — If you set this option, the database appends the USING parallelism = <value> clause to all intermediate SQL queries the model executes during training where value is the specified positive integer value. The default value is unspecified. In this case, the database does not add this clause.
ridgeCoefficient — (BETA. This option is part of an in-development feature and is subject to change.) If you specify this option, the value must be a nonnegative double. This option adds L2 (ridge) regularization to the normal equation by adding this value times the identity matrix to the X-transpose-X (Gram) matrix before inversion. Larger values shrink the coefficients more toward zero. When normalize=true (the default), the penalty is applied in standardized space and the persisted coefficients are un-scaled to the original units, matching scikit-learn’s Ridge(normalize=True) convention. The default value is 0.0 (no regularization). You can combine this option with the lassoCoefficient option for elastic net regularization.
suppressArrayLengthCheck — If you set this option, the featureArray option must be set to true. The system skips checking that the array length is the same size for all rows in the input. The default value is false.
threshold — This option enables soft thresholding. If you specify this option, then the option must be a positive numeric value. After the model calculates the coefficients, if any of them are greater than the threshold value, the threshold value is subtracted from them. If any coefficients are less than the negation of the threshold value, the model adds the threshold value to them. For any coefficients that are between the negative and positive threshold values, the model sets those coefficients to zero. Last modified on October 5, 2026