Stratified K-Fold iterator variant with non-overlapping groups.
This cross-validation object is a variation of StratifiedKFold attempts to return stratified folds with non-overlapping groups. The folds are made by preserving the percentage of samples for each class.
Each group will appear exactly once in the test set across all folds (the number of distinct groups has to be at least equal to the number of folds).
The difference between GroupKFold and StratifiedGroupKFold is that the former attempts to create balanced folds such that the number of distinct groups is approximately the same in each fold, whereas StratifiedGroupKFold attempts to create folds which preserve the percentage of samples for each class as much as possible given the constraint of non-overlapping groups between splits.
Read more in the User Guide.
For visualisation of cross-validation behaviour and comparison between common scikit-learn split methods refer to Visualizing cross-validation behavior in scikit-learn
Number of folds. Must be at least 2.
Whether to shuffle each class’s samples before splitting into batches. Note that the samples within each split will not be shuffled. This implementation can only shuffle groups that have approximately the same y distribution, no global shuffle will be performed.
When shuffle is True, random_state affects the ordering of the indices, which controls the randomness of each fold for each class. Otherwise, leave random_state as None. Pass an int for reproducible output across multiple function calls. See Glossary.
See also
StratifiedKFoldTakes class information into account to build folds which retain class distributions (for binary or multiclass classification tasks).
GroupKFoldK-fold iterator variant with non-overlapping groups.
The implementation is designed to:
y = ["Happy", "Sad"] to y = [1, 0] should not change the indices generated.>>> import numpy as np
>>> from sklearn.model_selection import StratifiedGroupKFold
>>> X = np.ones((17, 2))
>>> y = np.array([0, 0, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0])
>>> groups = np.array([1, 1, 2, 2, 3, 3, 3, 4, 5, 5, 5, 5, 6, 6, 7, 8, 8])
>>> sgkf = StratifiedGroupKFold(n_splits=3)
>>> sgkf.get_n_splits(X, y)
3
>>> print(sgkf)
StratifiedGroupKFold(n_splits=3, random_state=None, shuffle=False)
>>> for i, (train_index, test_index) in enumerate(sgkf.split(X, y, groups)):
... print(f"Fold {i}:")
... print(f" Train: index={train_index}")
... print(f" group={groups[train_index]}")
... print(f" Test: index={test_index}")
... print(f" group={groups[test_index]}")
Fold 0:
Train: index=[ 0 1 2 3 7 8 9 10 11 15 16]
group=[1 1 2 2 4 5 5 5 5 8 8]
Test: index=[ 4 5 6 12 13 14]
group=[3 3 3 6 6 7]
Fold 1:
Train: index=[ 4 5 6 7 8 9 10 11 12 13 14]
group=[3 3 3 4 5 5 5 5 6 6 7]
Test: index=[ 0 1 2 3 15 16]
group=[1 1 2 2 8 8]
Fold 2:
Train: index=[ 0 1 2 3 4 5 6 12 13 14 15 16]
group=[1 1 2 2 3 3 3 6 6 7 8 8]
Test: index=[ 7 8 9 10 11]
group=[4 5 5 5 5]
Get metadata routing of this object.
Please check User Guide on how the routing mechanism works.
A MetadataRequest encapsulating routing information.
Returns the number of splitting iterations in the cross-validator.
Always ignored, exists for compatibility.
Always ignored, exists for compatibility.
Always ignored, exists for compatibility.
Returns the number of splitting iterations in the cross-validator.
Request metadata passed to the split method.
Note that this method is only relevant if enable_metadata_routing=True (see sklearn.set_config). Please see User Guide on how the routing mechanism works.
The options for each parameter are:
True: metadata is requested, and passed to split if provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it to split.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.
Added in version 1.3.
Note
This method is only relevant if this estimator is used as a sub-estimator of a meta-estimator, e.g. used inside a Pipeline. Otherwise it has no effect.
Metadata routing for groups parameter in split.
The updated object.
Generate indices to split data into training and test set.
Training data, where n_samples is the number of samples and n_features is the number of features.
The target variable for supervised learning problems.
Group labels for the samples used while splitting the dataset into train/test set.
The training set indices for that split.
The testing set indices for that split.
© 2007–2025 The scikit-learn developers
Licensed under the 3-clause BSD License.
https://scikit-learn.org/1.6/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html