Ziyan Xia
Ziyan (Cecilia) Xia

Ziyan Xia

Technology · Art · Design

Artsy by instinct.
Data science by degree.

Building at the intersection of technology, art, and design — product analytics and insight, and storytelling through art direction.

Cecilia · San Francisco Bay Area · 夏子言

Experience

Data Analyst / Analytics Engineer

Revvo Technologies Inc. · AI road safety platform

Jul 2022 – Dec 2025 · San Mateo, CA

  • Metric design and product insights. Built end-to-end KPI frameworks for a large-scale IoT ecosystem (1B+ miles, 5K+ tire events), defining self-serve real-time metrics and customer-facing insights that enabled proactive road-safety alerts and contributed $75M+ in customer operating cost savings.
  • Forecasting and machine learning. Deployed production regression, classification, clustering, and survival models for tire risk prediction, sensor positioning, and depot detection — up to 90% accuracy across 50+ enterprise fleets in 25+ U.S. states.
  • Experimentation and causal inference. Designed and ran field A/B tests evaluating firmware updates, gateway repositioning, and signal-filtering algorithms, comparing telemetry distributions and downstream model outputs to prevent production regressions at scale.
  • Anomaly detection and root cause analysis. Automated investigation of abnormal sensor behavior with Python/SQL diagnostics and cross-device comparisons, separating environmental noise from hardware faults and cutting false-positive tire replacements by 80%+ (200+/week to under 40/week).
  • Production analytics workflows. Engineered analytics pipelines (Python, SQL/GCP, Java, Looker) behind Revvo's TireIQ platform, with Git, unit testing, and CI/CD for reliable deploys — investigation time down 75% (2 hrs to 30 min).
  • Algorithm design and optimization. Designed sensor auto-registration algorithms using statistical methods and GPS telemetry, correcting noisy mappings and false positives and reducing installation calibration time 85% (5 days to 18 hrs).
  • Cross-functional work. Partnered with Product, Engineering, and GTM teams to define success metrics, evaluate feature launches and account health, and deliver recommendations guiding product iteration, ML model calibration, LLM agent evaluation, and C-suite decisions.
  • IoT
  • B2B SaaS
  • Experimentation
  • Forecasting
  • Looker
  • GTM Partnership

Undergraduate Research Assistant

University of California, San Francisco · Roland Henry Lab
through the UC Berkeley URAP program

Feb 2020 – Feb 2021 · San Francisco, CA

  • Built Python analytics and modeling pipelines on high-frequency wearable (Fitbit) clinical data, using ARIMA-regression hybrid models to forecast mobility trends across longitudinal patient cohorts.
  • Established subject-level baselines and deviation signals through within-subject analysis and hypothesis testing, improving detection of meaningful motor-function changes over time.
  • Cleaned and transformed longitudinal clinical data from disparate sources — missing-value imputation, noise reduction, validation rules, and feature extraction — to create stable inputs for downstream modeling.
  • Ran exploratory analysis of clinical time-series data to surface behavioral patterns, seasonal effects, and activity cycles, supporting hypothesis generation and model design.
  • Time Series
  • Wearables
  • Clinical Data
  • ARIMA

Education

M.S., Statistics

Carnegie Mellon University

Aug 2021 – May 2022 · Pittsburgh, PA · GPA 3.95 / 4.0

B.S., Statistics

Central China Normal University

Sep 2017 – Jun 2021 · Wuhan, China · GPA 87.69 / 100

Visiting Student (BISP), Statistics

University of California, Berkeley

Aug 2019 – May 2020 · Berkeley, CA · GPA 3.5 / 4.0

Skills

  • Analytics & experimentation. A/B testing, experiment design, causal inference, metric design, time series, LLM, NLP
  • Programming. Python (NumPy, pandas, scikit-learn, statsmodels, TensorFlow, PyTorch), SQL, Java, R, JavaScript, Bash
  • Data engineering. BigQuery/PostgreSQL, Spark, Airflow, ETL pipelines, batch processing, data modeling, API integration
  • Visualization. Looker, Tableau, Plotly, Power BI, R Shiny, Excel/Google Sheets
  • Production practices. Git, Docker, CI/CD, unit testing, Agile/Scrum (Jira)

Projects

Bike Share Demand & Rebalancing

Predictive modeling · Washington, D.C.

  • Predictive model for bike availability, used to inform reshuffling strategy across the city from multi-year usage data.

Fitbit peak detection & DTW clustering

Time series · R

  • Peak detection on Fitbit activity data, with dynamic time warping used to cluster daily activity patterns.

Dockerized Python functions for PostgreSQL

Data engineering · Python

  • Walkthrough of deploying Python functions for PostgreSQL inside Docker containers.

XGBoost vs. iterative Random Forest on imbalanced high-dimensional genetic data

Write-up (PDF)

  • Comparison on imbalanced, high-dimensional genetic datasets; iterative Random Forest outperformed XGBoost.

About Me

I am a product data scientist who builds end-to-end analytics and modeling systems for B2B SaaS products. Most recently I designed production pipelines and dashboards that cut decision latency by 32% and contributed over $75M in customer cost savings through tire-health insights.

I like the part of the work that sits between teams — partnering with product, engineering, and go-to-market to turn large-scale data into decisions people actually act on.

Certifications

Social and Behavioral Research — Basic/Refresher

CITI Program

Issued Feb 2022 · Expired Feb 2025 · ID 47150531

Volunteering

English Teacher Volunteer

Arsasom Organization (Thailand NGO)

Jan 2019

  • Taught English and created instructional media at Phasukmaneejakmittraphab 116 School, a public school in Thailand.

Elsewhere