Data Engineering and Analytics for Insurance Risk Assessment
Pipelines and analytics for insurance risk scoring and fraud detection.
Key Details
| Challenge | Risk and fraud teams waited on manual extracts; models could not see fresh features. |
|---|---|
| Solution | A warehouse, feature jobs and dashboards for risk scoring and investigation queues. |
| Technologies | dbt, Snowflake, Python, Power BI |
Technologies used
Client background
An insurance carrier’s risk and fraud teams depended on weekly spreadsheet extracts. Models trained on stale features, investigators chased cases without a shared queue, and audits asked for lineage nobody could produce.
Key challenges
- Manual extracts delayed risk refresh and created untracked copies of sensitive data.
- Feature definitions differed between data science notebooks and production jobs.
- Fraud investigators lacked a prioritized queue tied to model scores.
- No reliable audit trail from source system to dashboard metric.
What we built
- Snowflake warehouse with dbt models for claims, policy and behavioral features.
- Scheduled feature jobs with tests and documentation for model consumers.
- Power BI risk and fraud dashboards with investigation queues.
- Lineage and access patterns designed for audit and least-privilege use.
Project team: 8 engineers across AI/ML, backend and domain specialists — delivery over 18 weeks.
How we delivered
01
Map
Traced source systems, SLAs and which metrics risk leadership actually used.
02
Model
Built a star schema and feature contracts shared by analytics and ML.
03
Pipeline
Automated daily refresh with tests for freshness and null rates.
04
Enable
Trained investigators and risk ops on the new dashboards and queues.
Business impact
DailyRisk refresh
FewerManual extracts
AuditTrail on features