Senior Data Engineer

Truthset
Truthset

Data Science

United States

Posted on Oct 4, 2026
Senior Data Engineer

Level: Senior (6+ years of experience)

Location: New York, or Remote (US)

Who We Are

Truthset is a venture-backed data quality and identity company solving the multi-billion dollar problem of data accuracy for the marketing industry. We work with a collective of leading data providers to measure and validate the accuracy of consumer data, and we turn that into products — data-rated audiences, accuracy scores, and identity resolution — used by brands, agencies, publishers, and platforms such as Paramount, Procter & Gamble, and TransUnion to improve precision and marketing ROI. We are a small, senior team that moves fast and owns its work end to end.

Our Tech Stack: Databricks (Delta Lake, Unity Catalog, Delta Sharing, Workflows), Spark (PySpark / Spark SQL), Python, SQL, PostgreSQL, AWS (S3, EC2, IAM), Airflow, Git/CI-CD.

About the Role

We are looking for a Senior Data Engineer to own and scale the data platform behind our accuracy, audience, and identity products. You will be responsible for a growing portfolio of pipelines that bring in data from our data providers and partners, for the privacy compliance and opt-out pipelines that keep that data lawful, and for making our processing fast and cost-efficient at billion-record scale. A major part of the role is leading the migration of performance-critical, application-level PostgreSQL workloads into Databricks. This is a hands-on, high-ownership role with real influence on architecture and engineering standards. Your experience and intuition in designing robust and reusable workflows, managing data lifecycles, and thinking in the big picture about data governance organization-wide, plus your motivation to be a key owner and invested decision maker, will be invaluable in taking Truthset’s data platform to the next level.

What You'll Do
  • Own provider and partner data pipelines — onboard new data providers and partners, and operate a growing set of ingestion pipelines with clear schemas, data contracts, validation, and SLAs.
  • Build and run compliance and opt-out pipelines — design auditable pipelines that process consumer deletion, opt-out, and suppression requests (e.g., California Delete Act/DROP, CCPA/CPRA, other state privacy laws) and propagate them reliably across raw, derived, and delivered datasets, including verifiable physical purge in Delta Lake and cloud storage.
  • Migrate PostgreSQL workloads to Databricks — lead the re-architecture of application-level Postgres functionality (e.g., identity resolution, matching, and reporting) into Spark/Delta, designed for distributed execution rather than lift-and-shift, with parity testing and safe cutover.
  • Optimize performance and cost at scale — tune Spark and Delta workloads over billions of rows: data layout (partitioning, liquid clustering, Z-ordering), join strategy, skew, file sizing, caching, and cluster/serverless configuration.
  • Model data for our products — design lakehouse data models (medallion architecture) that support accuracy scoring, audience generation, identity graphs, and provider attribution.
  • Govern and share data — manage access, lineage, and cataloging in Unity Catalog, and support secure data exchange with partners via Delta Sharing and other delivery mechanisms.
  • Ensure data quality and reliability — implement automated quality checks, monitoring, and alerting; lead troubleshooting and root-cause analysis for data incidents.
  • Support outbound deliveries — build and maintain reliable data feeds to activation and marketplace partners.
  • Engineer for production — write clean, tested, well-documented Python and SQL; deploy through CI/CD and infrastructure-as-code (e.g., Databricks Asset Bundles, Terraform).
  • Lead technically — shape architecture decisions, set engineering standards, review code, mentor others, and partner closely with data science, product, and data operations.
What You'll Bring
  • 6+ years of experience in data engineering or a closely related role, including ownership of production pipelines at large scale.
  • 2+ years of hands-on production experience with Databricks.
  • Expert-level Spark (PySpark and Spark SQL), Python, and SQL.
  • Deep working knowledge of Delta Lake — MERGE/upserts, change data feed, schema evolution, time travel, OPTIMIZE/VACUUM, and deletion vectors.
  • A track record of diagnosing and fixing performance problems on datasets with hundreds of millions to billions of rows.
  • Strong PostgreSQL skills (query plans, indexing, tuning) and experience migrating relational/application workloads to a distributed platform.
  • Experience ingesting and operating many external data feeds from third parties with varied formats, delivery methods (S3, SFTP, data sharing), and quality levels.
  • Hands-on experience implementing privacy compliance in data systems — deletion, suppression, retention, and auditability under CCPA/CPRA, GDPR, or similar.
  • Solid AWS experience (S3, IAM, multi-account environments) and orchestration experience (Databricks Workflows/Jobs, Airflow, or similar).
  • Strong Git, testing, and CI/CD practices.
  • Clear communication, sound judgment, and comfort owning outcomes in a small team where priorities evolve quickly.
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
Nice to Have
  • Identity resolution experience — hashed emails (HEMs), device IDs, identity graphs, and entity resolution/matching at scale.
  • Ad-tech or mar-tech background — audience activation through DSPs/SSPs and data marketplaces, measurement and impression-log data, or data clean rooms.
  • Unity Catalog administration, Delta Sharing, Lakeflow/Delta Live Tables, Databricks Asset Bundles, or Terraform.
  • Experience with other cloud data platforms, such as Snowflake.
  • Databricks Certified Data Engineer Professional (or equivalent).
  • Experience using AI-assisted development tools effectively in a production engineering workflow.
  • Experience at an early-stage or growth-stage startup.
How to Apply

Send your resume and a short note about why you're interested to careers@truthset.com.