Machine Learning Models are only as Good as the Data They’re Trained on

This blog post was written by Inderjit Singh Chahal as part of the Deepchecks Community Blog. If you would like to contribute your own blog post, feel free to reach out to us via blog@deepchecks.com. We typically pay a symbolic fee for content that's accepted by our reviewers.

Introduction

The quality of a Machine Learning model is decided largely by the quality of the dataset it was trained on. A recent study compiled responses from a number of Machine Learning practitioners highlighted that access, preparation, and validation of data  are the most time-consuming components in a Machine Learning project. The time distribution in that study indicates the percentage  of time taken by different components of a Machine Learning production cycle:

Furthering the argument on the importance of data, the majority of the respondents in that study revealed that data collection, preparation, and validation are the most critical components in their projects. This figure outlines those responses:

Data Validation Techniques and Tools

Now that we have established the importance,let us discuss the available tools and techniques. These techniques are generally classified into two broader categories:

  • Proactive Data Validation
  • Reactive Data Validation

As the name suggests, we prefer to use Proactive Data Validation in most practical implementations since it saves us plenty of time by taking care of issues in the earlier modeling steps. It is almost always the most critical component associated with data validation. The primary reason being its association with the gold standard dataset that forms the bedrock for all downstream decisions such as:

  • Feature Selection
  • Feature Importance Calculation
  • Model Selection
  • Model Validation
  • Benchmarking

Proactive Data Validation Tools and Techniques

Proactive Data Validation tools and techniques are further classified into:

– Type Safety is when we have a middleware that validates the data types and other implementations with integration to the main annotations tool (at source) to prevent errors in the remainder of the downstream tasks.

Testing. CI/CD. Monitoring.

Because ML systems are more fragile than you think. All based on our open-source core.

Our GithubInstall Open SourceBook a Demo

Tool to Solve This

Below is a small snippet that displays a  type safety constraint :

def wrapper(func):

    def inner(foo,bar):
        for each_arg in func.__code__.co_varnames:
            if not type(eval(each_arg)) == func.__annotations__[each_arg]:
                raise TypeError(f"expected dtype {func.__annotations__[each_arg]} but got {type(eval(each_arg))}")

        return func(foo,bar)

    return inner

@wrapper
def middleware_func(foo: int, bar: str) -> (str,int):
    return "out"

– Schema Validation is when we are validating the annotated data on the storage side where we expect to have integers (e.g., coordinates of bounding boxes should be digits) or other data types where we can enforce these data types.

import schemathesis

schema = schemathesis.from_uri("http://example.com/swagger.json")


@schema.parametrize()
def test_api(case):
    case.call_and_validate()

– Label Ambiguity Search looks for identical samples with different labels. This is likely caused by mislabeled data or when data collection has missing features.

The Tool to Solve This
Below is a sample snippet that can be used as one of the automated data validation tools, using deepchecks library to search for data ambiguity in a data verification pipeline:

from deepchecks.checks.integrity import LabelAmbiguity
from deepchecks.base import Dataset
import pandas as pd

check = LabelAmbiguity()
check.add_condition_ambiguous_sample_ratio_not_greater_than(0)
result = check.run(phishing_dataset)
result.show(show_additional_outputs=False

# Output

– A/B Testing in Data Validation for Machine Learning is the controlled experiments used to validate the analytics flow against a hold-out gold standard data set where we know with 100% certainty it is accurate and represents the production dataset.

Reactive Data Validation Techniques and Tools

– The Freshness Testing technique is used to determine how up-to-date the data sources are for the retraining pipelines of a model.

The Tools to Solve This

We can use tools like dbt, Tensorflow data validation to validate the health and relevance of the data on some set preconditions to ensure that the data quality validation is thoroughly completed and we are using the most relevant data for production training/fine-tuning pipelines.

–  Distribution/Drift Checks allows us to keep track of the changes in the distribution of the incoming data (test or inference data) that might affect a model’s relevance to make predictions.

The Tool to Solve This

Deepchecks provides a very simple way of keeping track of these shifts. The small code snippet below ensures data quality validation and keeps track of any changes in the data distribution:

check  = TrainTestDrift()
result  = check.run(train_dataset=train_dataset, test_dataset=test_dataset, model=model)

Source: Deepchecks

– Feedback Loop for Skew Monitoring is used to monitor skews that occur based on how the prediction results are presented to the user. For example, if a recommender system is scoring 100 videos for a user and the user is only presented with the top 10 results, the remaining 90 is not going to receive any attention and will never be part of the feedback loop.

The Tool to Solve This

We can infer from the definition that this is more of a data pipeline issue, which is why we should visualize the pipelines using Graphviz to ensure we have adequate checks and that we don’t have this skew in our production pipeline. Below is a simple graph generated for such a pipeline using Graphviz:

Conclusion

We established the importance of data validation in Machine Learning and looked into the various available tools for different sets of data validation techniques and procedures. A data validation tool can be a sophisticated UI/UX or a small code snippet that keeps track of the changes that are happening in the environment of the model. These tools not only make lives easier by automating most of the mundane tasks associated with the data validation process, but are critical in the maintenance of Machine Learning pipelines in production.

Testing. CI/CD. Monitoring.

Because ML systems are more fragile than you think. All based on our open-source core.

Our GithubInstall Open SourceBook a Demo
×
Deepchecks is joining forces with Check Point Strengthening AI security – together.