← Back to jobs

Senior Data Analyst (Graph focused), Fixed Term

Location
Toronto, CA
Work type
Full Time
Posted
2026-07-20

Job description

The Role:

The Senior Data Analyst is a hands-on analytical contributor embedded in the engineering team. This role bridges the gap between the data itself and the engineers, product managers, and data analysts building and evaluating the Knowledge Graph. The primary focus is data integrity: validating datasets, verifying query outputs, tracing the root cause of discrepancies, and applying statistical methods to assess data quality across multiple storage technologies.
This is a practitioner role, not a consulting engagement. The deliverable is evidence — validated results, documented defects, root cause analysis, and statistical assessments that the team can act on.

What You’ll Do:
Validate datasets loaded into each technology to confirm completeness, accuracy, and structural integrity relative to source data in Databricks and CDS
Verify benchmark query outputs across technologies — confirm that the same logical query against the same underlying data produces consistent, correct results regardless of which system executes it
Identify, document, and trace the root cause of data discrepancies and defects discovered during validation; distinguish between ETL issues, schema translation errors, technology-specific behavior, and upstream data quality problems
Develop and maintain validation test cases and expected outputs for benchmark queries and compliance use cases
Support the collection and documentation of compliance use cases from the business unit
Help assess each use case: can it be fulfilled by a conventional database approach, or does it require the traversal and pattern-matching capabilities of a purpose-built graph database?
Contribute analytical rigor to use case triage — this is a cost-and-complexity decision as much as a technical one
Apply statistical methods to evaluate dataset representativeness, sampling quality, and measurement reliability across benchmark runs
Analyze benchmark result distributions — identify outliers, assess variance across cold/warm/concurrent runs, and flag results that require deeper investigation before scoring
Produce summary statistics and data quality reports that inform the team’s architecture assessment
Document validation findings, defect reports, and root cause analyses in a format the engineering team can act on
Maintain a running record of known data issues and their resolution status across each technology under evaluation

What You Bring:

Demonstrated experience validating large, complex datasets — identifying discrepancies, tracing root causes, and documenting findings clearly
Strong SQL skills; ability to write analytical queries against relational databases (PostgreSQL experience preferred)
Experience working with data at significant scale — hundreds of millions of records — where manual spot-checking is insufficient and systematic validation approaches are required
Familiarity with ETL pipelines and the types of data quality issues that arise in data loading and transformation
Practical experience applying statistical methods to data quality assessment: distribution analysis, outlier detection, variance analysis, sampling validation
Ability to interpret benchmark result data and distinguish meaningful performance differences from noise
Comfort working across multiple database technologies and query languages — this role will need to query data in PostgreSQL, graph databases, and Databricks as part of normal validation work
Experience with Databricks or similar distributed data platforms (Spark, Delta Lake)
Strong written communication — validation findings, defect reports, and root cause analyses must be clear enough for both engineers and product stakeholders
Ability to work independently under minimal supervision, taking direction from peers rather than requiring structured management oversight
Experience embedded in a cross-functional engineering team

Original source