w.feature_engineering: Feature Engineering

class databricks.sdk.service.ml.FeatureEngineeringAPI

[description]

backfill_features(feature_full_names: List[str], backfill_ranges: List[BackfillRange] [, budget_policy_id: Optional[str], request_id: Optional[str], tags: Optional[Dict[str, str]]]) → BackfillFeaturesOperation

Backfill features.

Parameters:
  • feature_full_names – List[str] Full names of the features to backfill.

  • backfill_ranges – List[BackfillRange] Output ranges to backfill.

  • budget_policy_id – str (optional) The budget policy ID, in UUID format, used to attribute the serverless compute cost of this backfill. If not specified, a default budget policy may be applied.

  • request_id – str (optional) Idempotency token for the request.

  • tags – Dict[str,str] (optional) Custom tags to associate with this backfill. They are applied to the backfill job and forwarded to the underlying compute as Databricks resource tags, so backfill cost can be attributed in the billing system tables. These tags apply only to the backfill compute; they are not applied to the Unity Catalog Feature resources themselves, whose tags are managed separately through the Unity Catalog tagging API. A maximum of 25 tags is supported; keys and values are subject to the same limitations as Databricks resource tags.

Returns:

Operation

batch_create_materialized_features(requests: List[CreateMaterializedFeatureRequest]) → BatchCreateMaterializedFeaturesResponse

Batch create materialized features.

Parameters:

requests – List[CreateMaterializedFeatureRequest] The requests to create materialized features.

Returns:

BatchCreateMaterializedFeaturesResponse

cancel_operation(name: str)

Cancel an operation.

Parameters:

name – str The name of the operation resource to be cancelled.

create_feature(feature: Feature) → Feature

Create a Feature.

Parameters:

feature – Feature Feature to create.

Returns:

Feature

create_kafka_config(kafka_config: KafkaConfig) → KafkaConfig

Create a Kafka config. During PrPr, Kafka configs can be read and used when creating features under the entire metastore. Only the creator of the Kafka config can delete it.

Parameters:

kafka_config – KafkaConfig

Returns:

KafkaConfig

create_materialized_feature(materialized_feature: MaterializedFeature) → MaterializedFeature

Create a materialized feature.

Parameters:

materialized_feature – MaterializedFeature The materialized feature to create.

Returns:

MaterializedFeature

create_stream(stream: Stream) → Stream

Create a Stream, a governed UC entity representing an external streaming data source.

Parameters:

stream – Stream The Stream to create.

Returns:

Stream

delete_feature(full_name: str)

Delete a Feature.

Parameters:

full_name – str Name of the feature to delete.

delete_kafka_config(name: str)

Delete a Kafka config. During PrPr, Kafka configs can be read and used when creating features under the entire metastore. Only the creator of the Kafka config can delete it.

Parameters:

name – str Name of the Kafka config to delete.

delete_materialized_feature(materialized_feature_id: str)

Delete a materialized feature.

Parameters:

materialized_feature_id – str The ID of the materialized feature to delete.

delete_stream(name: str)

Delete a Stream by its full three-part name (catalog.schema.stream).

Parameters:

name – str Full three-part name (catalog.schema.stream) of the Stream to delete.

get_feature(full_name: str) → Feature

Get a Feature.

Parameters:

full_name – str Name of the feature to get.

Returns:

Feature

get_kafka_config(name: str) → KafkaConfig

Get a Kafka config. During PrPr, Kafka configs can be read and used when creating features under the entire metastore. Only the creator of the Kafka config can delete it.

Parameters:

name – str Name of the Kafka config to get.

Returns:

KafkaConfig

get_materialized_feature(materialized_feature_id: str) → MaterializedFeature

Get a materialized feature.

Parameters:

materialized_feature_id – str The ID of the materialized feature.

Returns:

MaterializedFeature

get_operation(name: str) → Operation

Get an operation.

Parameters:

name – str The name of the operation resource.

Returns:

Operation

get_stream(name: str) → Stream

Get a Stream by its full three-part name (catalog.schema.stream).

Parameters:

name – str Full three-part name (catalog.schema.stream) of the Stream to get.

Returns:

Stream

list_features(catalog_name: str, schema_name: str [, page_size: Optional[int], page_token: Optional[str]]) → Iterator[Feature]

List Features.

Parameters:
  • catalog_name – str Name of parent catalog for features of interest.

  • schema_name – str Name of parent schema relative to its parent catalog.

  • page_size – int (optional) The maximum number of results to return.

  • page_token – str (optional) Pagination token to go to the next page based on a previous query.

Returns:

Iterator over Feature

list_kafka_configs([, page_size: Optional[int], page_token: Optional[str]]) → Iterator[KafkaConfig]

List Kafka configs. During PrPr, Kafka configs can be read and used when creating features under the entire metastore. Only the creator of the Kafka config can delete it.

Parameters:
  • page_size – int (optional) The maximum number of results to return.

  • page_token – str (optional) Pagination token to go to the next page based on a previous query.

Returns:

Iterator over KafkaConfig

list_materialized_features([, feature_name: Optional[str], page_size: Optional[int], page_token: Optional[str]]) → Iterator[MaterializedFeature]

List materialized features.

Parameters:
  • feature_name – str (optional) Filter by feature name. If specified, only materialized features materialized from this feature will be returned.

  • page_size – int (optional) The maximum number of results to return. Defaults to 100 if not specified. Cannot be greater than 1000.

  • page_token – str (optional) Pagination token to go to the next page based on a previous query.

Returns:

Iterator over MaterializedFeature

list_streams([, page_size: Optional[int], page_token: Optional[str], parent: Optional[str]]) → Iterator[Stream]

List Streams under a given catalog.schema parent.

Parameters:
  • page_size – int (optional) The maximum number of results to return.

  • page_token – str (optional) Pagination token to go to the next page based on a previous query.

  • parent – str (optional) Two-part name (catalog.schema) of the parent under which to list Streams.

Returns:

Iterator over Stream

purge_feature_entities(features: List[str], entities_table: str [, budget_policy_id: Optional[str], request_id: Optional[str], tags: Optional[Dict[str, str]]]) → PurgeFeatureEntitiesOperation

Purge materialized feature values for specified entities.

Parameters:
  • features – List[str] Fully qualified names of the features to purge. At least one nonempty feature name is required. A request may contain at most 10000 features; submit additional features in separate requests. Duplicate features are rejected.

  • entities_table – str Fully qualified name of the Unity Catalog Delta table containing the entity keys to purge. The table may contain a subset of each feature’s entity-key columns. A partial key match deletes all feature rows matching the provided key values. Non-key columns are rejected; null key values are allowed.

  • budget_policy_id – str (optional) The budget policy ID, in UUID format, used to attribute the serverless compute cost of this purge. If not specified, a default budget policy may be applied.

  • request_id – str (optional) Optional UUID4 idempotency token for the request.

  • tags – Dict[str,str] (optional) Custom tags to associate with this purge. They are applied to the purge job and forwarded to the underlying compute as Databricks resource tags, so purge cost can be attributed in the billing system tables. These tags apply only to the purge compute; they are not applied to the Unity Catalog Feature resources themselves, whose tags are managed separately through the Unity Catalog tagging API. A maximum of 25 tags is supported; keys and values are subject to the same limitations as Databricks resource tags.

Returns:

Operation

update_feature(full_name: str, feature: Feature, update_mask: str) → Feature

Update a Feature.

Parameters:
  • full_name – str The full three-part name (catalog, schema, name) of the feature. This is the feature’s resource identifier; the catalog_name, schema_name, and name fields below are OUTPUT_ONLY decomposed views of this value.

  • feature – Feature Feature to update.

  • update_mask – str The list of fields to update.

Returns:

Feature

update_kafka_config(name: str, kafka_config: KafkaConfig, update_mask: FieldMask) → KafkaConfig

Update a Kafka config. During PrPr, Kafka configs can be read and used when creating features under the entire metastore. Only the creator of the Kafka config can delete it.

Parameters:
  • name – str Name that uniquely identifies this Kafka config within the metastore. This will be the identifier used from the Feature object to reference these configs for a feature. Can be distinct from topic name.

  • kafka_config – KafkaConfig The Kafka config to update.

  • update_mask – FieldMask The list of fields to update.

Returns:

KafkaConfig

update_materialized_feature(materialized_feature_id: str, materialized_feature: MaterializedFeature, update_mask: str) → MaterializedFeature

Update a materialized feature (pause/resume).

Parameters:
  • materialized_feature_id – str Server-assigned unique identifier for the materialized feature.

  • materialized_feature – MaterializedFeature The materialized feature to update.

  • update_mask – str Provide the materialization feature fields which should be updated. Currently, only the pipeline_state field can be updated.

Returns:

MaterializedFeature

update_stream(name: str, stream: Stream, update_mask: FieldMask) → Stream

Update a Stream. Only fields listed in update_mask are mutated.

Parameters:
  • name – str Full three-part (catalog.schema.stream) name of the stream.

  • stream – Stream The Stream to update.

  • update_mask – FieldMask The list of fields to update.

Returns:

Stream