最新的Databricks Certified Machine Learning Professional - Databricks-Machine-Learning-Professional免費考試真題
A machine learning engineer has developed a model and registered it using the FeatureStoreClient fs. The model has model URI model_uri. The engineer now needs to perform batch inference on customer-level Spark DataFrame spark_df, but it is missing a few of the static features that were used when training the model. The customer_id column is the primary key of spark_df and the training set used when training and logging the model.
Which of the following code blocks can be used to compute predictions for spark_df when the missing feature values can be found in the Feature Store by searching for features by customer_id?
Which of the following code blocks can be used to compute predictions for spark_df when the missing feature values can be found in the Feature Store by searching for features by customer_id?
正確答案: C
A data scientist is utilizing MLflow to track their machine learning experiments. After completing a run with run ID run_id for the experiment with experiment ID exp_id, the data scientist wants to programmatically return the logged metrics for run_id. They have an active MLflow Client client and an active Spark session spark. Which lines of code can be used to return the logged metrics for run_id?
正確答案: A
說明:(僅 Fast2test 成員可見)
A machine learning engineer is monitoring label values for a production machine learning classification model. The engineer believes that the relative prevalence of the classes is becoming changing in more recent data. Which tool can the machine learning engineer use to assess their theory?
正確答案: C
說明:(僅 Fast2test 成員可見)
A Machine Learning Engineer has a real-time fraud detection model deployed that approves or blocks millions of transactions daily. They need to deploy a new version of the model with improved detection accuracy to this high-traffic, business-critical application. Because any model downtime could result in lost revenue or customer dissatisfaction, the engineer must ensure zero downtime and minimal disruption for end users. Leadership also requires that any rollback to the previous version be immediate if issues are detected with the new model in production. Which deployment strategy meets these requirements?
正確答案: B
說明:(僅 Fast2test 成員可見)
Why are Delta tables often used to store machine learning features?
正確答案: D
說明:(僅 Fast2test 成員可見)
Which of the following is a drawback associated with using Jensen-Shannon (JS) distance for numeric feature drift detection?
正確答案: B
說明:(僅 Fast2test 成員可見)
A data scientist has written a function to track the runs of their random forest model. The data scientist is changing the number of trees in the forest across each run. Which of the following MLflow operations is designed to log single values like the number of trees in a random forest?
正確答案: A
A Machine Learning Engineer is conducting hyperparameter tuning for multiple XGBoost models using Ray Tune on Databricks. They want to integrate MLflow tracking to monitor their experiments and need to ensure proper authentication. The engineer has Ray 2.41 installed and wants to use both Ray Tune and MLflow together in their distributed tuning workflow. They have to configure Databricks to run the hyperparameter optimization with MLflow integration. Which set of configuration steps will do this?
正確答案: B
說明:(僅 Fast2test 成員可見)
A Data Scientist is using Spark ML to train a model for detecting fraudulent transactions (1:
fraudulent, 0: non-fraudulent). To maximize the model's overall AUC (ROC) score, the Data Scientist wants to perform hyperparameter tuning on a Random Forest classifier. Due to budget constraints, the tuning process must be completed efficiently to avoid excessive compute costs.
Which approach will fulfill their needs?
fraudulent, 0: non-fraudulent). To maximize the model's overall AUC (ROC) score, the Data Scientist wants to perform hyperparameter tuning on a Random Forest classifier. Due to budget constraints, the tuning process must be completed efficiently to avoid excessive compute costs.
Which approach will fulfill their needs?
正確答案: C
說明:(僅 Fast2test 成員可見)
A Data Scientist needs to analyze drift detection results from Databricks Lakehouse Monitoring.
The system has generated both profile metrics and drift metrics tables. The scientist needs to identify baseline drift in numerical features by comparing current data against a baseline from 6 months ago. Which combination of table columns and values indicates baseline drift in a numerical feature?
The system has generated both profile metrics and drift metrics tables. The scientist needs to identify baseline drift in numerical features by comparing current data against a baseline from 6 months ago. Which combination of table columns and values indicates baseline drift in a numerical feature?
正確答案: B
說明:(僅 Fast2test 成員可見)
A data scientist is building a model to predict which communication channel (Phone, SMS, Email, or Post) is most likely to be effective for a given customer. Which model type is suited to this task?
正確答案: D
說明:(僅 Fast2test 成員可見)
A machine learning engineer needs to deliver predictions of a machine learning model in real- time. However, the feature values needed for computing the predictions are available one week before the query time. Which feature is a benefit of using a batch serving deployment in this scenario rather than a real-time serving deployment where predictions are computed at query time?
正確答案: C
A data scientist has developed a model to predict whether or not it will rain using the expected temperature and expected cloud coverage. However, the proportion of days where it actually rains has increased dramatically from the proportion in the data on which the model was trained.
Which type of drift is present in the above scenario?
Which type of drift is present in the above scenario?
正確答案: D
說明:(僅 Fast2test 成員可見)