Senior Data Analyst (Graph)

DS STREAMGdańsk, Kraków, Warsaw, Wrocławnofluffjobszveřejněno 24. 08. 2026
Nutné:Node.jsDataSenior

Daily tasks

  • Validate datasets loaded into each technology to confirm completeness, accuracy, and structural integrity relative to source data in Databricks and CDS
  • Verify benchmark query outputs across technologies — confirm that the same logical query against the same underlying data produces consistent, correct results regardless of which system executes it
  • Identify, document, and trace the root cause of data discrepancies and defects discovered during validation; distinguish between ETL issues, schema translation errors, technology-specific behavior, and upstream data quality problems
  • Develop and maintain validation test cases and expected outputs for benchmark queries and compliance use cases Use Case Assessment Support
  • Support the collection and documentation of compliance use cases from the business unit
  • Help assess each use case: can it be fulfilled by a conventional database approach, or does it require the traversal and pattern-matching capabilities of a purpose-built graph database
  • Contribute analytical rigor to use case triage — this is a cost-and-complexity decision as much as a technical one

Requirements

The Senior Data Analyst is a hands-on analytical contributor embedded in the engineering team. This role bridges the gap between the data itself and the engineers, product managers, and data analysts building and evaluating the Knowledge Graph. The primary focus is data integrity: validating datasets, verifying query outputs, tracing the root cause of discrepancies, and applying statistical methods to assess data quality across multiple storage technologies. This is a practitioner role, not a consulting engagement. The deliverable is evidence — validated results, documented defects, root cause analysis, and statistical assessments that the team can act on. Practical experience applying statistical methods to data quality assessment: distribution analysis, outlier detection, variance analysis, sampling validation Ability to interpret benchmark result data and distinguish meaningful performance differences from noise Cross-Technology Proficiency Comfort working across multiple database technologies and query languages — this role will need to query data in PostgreSQL, graph databases, and Databricks as part of normal validation work Experience with Databricks or similar distributed data platforms (Spark, Delta Lake) Communication and Collaboration Strong written communication — validation findings, defect reports, and root cause analyses must be clear enough for both engineers and product stakeholders Ability to work independently under minimal supervision, taking direction from peers rather than requiring structured management oversight Experience embedded in a cross-functional engineering team Strongly Preferred Qualifications Hands-on experience with graph databases (Neo4j, TigerGraph, or similar) — even at proof-of-concept scale Familiarity with graph data models: property graphs, node/edge schema, relationship taxonomies Familiarity with knowledge graph or ontology concepts

Must have: TigerGraph, Neo4j