Deepchecks Powers Evaluation for a Global Pharma’s Internal AI Platform

Customer Overview

A global pharmaceutical company leverages advanced AI to drive research and innovation. Its data and AI teams constantly seek technologies that enhance information reliability and streamline internal workflows across research teams.

The Challenge

The company’s internal platform is designed to assist research teams by parsing scientific PDFs, extracting structured information, and facilitating precise, AI-powered queries. However, the platform faced several challenges:

  • Limited visibility into the accuracy and robustness of AI-generated outputs. The lack of well-defined measurable results was a blocker both for validating system improvements during development, and for user adoption in production, due to hindered trust.
  • Problems with pipeline quality – such as inconsistent parsing and inaccurate data extraction from complex scientific PDFs, leading to data inaccuracies.

The Solution

The company’s innovation and AI teams collaborated with Deepchecks to address these challenges:

  • Integration of Deepchecks’ LLM Evaluation Platform: By integrating Deepchecks’ application, The team systematically evaluated the platform’s performance. The application enabled them to:
    • Track the accuracy and consistency of AI-generated answers.
    • Compare different pipeline versions across key evaluation metrics such as hallucination rate and factual accuracy.
    • Continuously monitor regressions and improvements through a centralized, visual dashboard.
  • Visibility to Evaluation Definition and Results:The team deployed a customizable evaluation widget, allowing users to filter by specific indications and tailored queries, improving user flexibility and quality of results, without additional engineering overhead.
  • Efficient enhancement to their PDF Parsing Pipeline:Utilizing Deepchecks’ automatic evaluation and root cause analysis, The company could immediately validate the progress of their parsing pipeline, significantly improving its coverage and the quality of the extracted information from complex scientific documents.

The Results

The collaboration yielded significant improvements:

  • 50% Faster Time to Production: By integrating Deepchecks’ evaluation workflows and systematic monitoring, The company reduced their time from development to production by approximately 50%, accelerating the delivery of their platform to internal users.
  • Significant Increase in Pipeline Quality: Improvements to the parsing and data extraction pipeline, coupled with Deepchecks’ monitoring, increased structured data extraction accuracy by 30%, boosting the reliability of downstream AI-generated answers.
  • Greater Trust and Ongoing Improvement: With continuous evaluation metrics like hallucination rate, factual accuracy, and extraction quality tracked over time, The company gained significantly greater visibility into system performance, enabling regular, targeted improvements and raising internal trust in their platform.

Through the collaboration with Deepchecks, The team both efficiently enhanced the performance and reliability of their platform and also established a robust evaluation framework for future innovation. The success of this initiative highlights their leadership in adopting cutting-edge AI tools to drive better research outcomes, while reinforcing Deepchecks’ role as a trusted partner for operationalizing and scaling GenAI applications.

Recent Case Studies

Enhancing Moovit’s Internal GenAI Pipeline with Deepchecks

Read Full Case Study

Lovehoney Group’s Journey to AI-Enhanced Customer Service with Deepchecks

Read Full Case Study
×
Deepchecks is joining forces with Check Point Strengthening AI security – together.