← All projects

CASE STUDY / 02

Energy AI Assistant Evaluation Platform

A repeatable evaluation workflow for measuring the quality, safety, unsupported certainty, and response speed of a bilingual energy-service AI assistant.

PythonpandasRegexTableau

01 / PROBLEM

Problem

Energy-service assistants answer questions about billing, pricing, outages, meters, safety, and conservation. A useful evaluation process must detect incomplete or unsafe answers, unsupported claims, language-level gaps, and slow responses across business categories.

02 / SOLUTION

What I Built

I built a structured evaluation pipeline for 20 simulated bilingual test cases across seven categories. It scores content completeness, checks safety and unsupported certainty, classifies failure types, calculates performance summaries, and feeds a Tableau dashboard for category and language analysis.

03 / PROCESS

Approach / Tools

  1. 01Designed bilingual test cases across seven energy-service categories with expected content and safety requirements.
  2. 02Evaluated response completeness, safety, unsupported certainty, and latency using pandas and regular-expression checks.
  3. 03Classified failure reasons and produced category-, language-, and KPI-level summary files for analysis.
  4. 04Built a Tableau dashboard to compare pass rates, failure patterns, hallucination rate, and response time.

04 / IMPACT

Key Results

20bilingual test cases
70.0%overall pass rate
10.0%hallucination rate
1.85 secaverage response time

05 / OUTPUT

Visualization

Tableau dashboard for the Energy AI Assistant Evaluation project
Tableau dashboard comparing headline KPIs, category performance, language performance, failures, and latency.

06 / SOURCE

GitHub Repository

Review the code, data workflow, reports, and project documentation.

View on GitHub ↗
Next case studyIndustrial Sales Analytics →