Senior Data Engineer

Posted 18 Days Ago
Hiring Remotely in USA
Remote
190K-220K Annually
Senior level
Software • Database
The Role
The Senior Data Engineer will build data ingestion and processing infrastructure using tools like Spark and AWS. Responsibilities include creating a data resolution framework, developing CI/CD pipelines, solving complex data problems, and collaborating with engineering and product teams on data-related issues.
Summary Generated by Built In

About Us

At People Data Labs, we’re committed to democratizing access to high-quality B2B data and leading the emerging DaaS economy. We empower developers, engineers, and data scientists to create innovative, compliant data products at scale with our clean, easy-to-use datasets of resume, company, location, and education data consumed through our suite of APIs. 

PDL is an innovative, fast-growing, global team backed by world-class investors, including Craft Ventures, Flex Capital, and Founders Fund. We scour the world for people hungry to improve, curious about how things work, and willing to challenge the status quo to build something new and better.

Roles & Responsibilities:

  • Build infrastructure for ingestion, transformation, and loading an exponentially increasing volume of data from a variety of sources using Spark, SQL, AWS, and Databricks
  • Building an organic entity resolution framework capable of correctly merging hundreds of billions of individual entities into a number of clean, consumable datasets.
  • Developing CI/CD pipelines and anomaly detection systems capable of continuously improving the quality of data we're pushing into production.
  • Devising solutions to largely-undefined data engineering and data science problems.
  • Work with stakeholders in Engineering and Product to assist with data-related technical issues and support their infrastructure needs

Technical Requirements

  • 5-7+ years industry experience with clear examples of strategic technical problem solving and implementation
  • Strong software development fundamentals.
  • Experience with Python 
  • Expertise with Apache Spark (Java, Scala, and/or Python-based)
  • Experience with SQL
  • Experience building scalable data processing systems (e.g., cleaning, transformation)  from the ground up.
  • Experience using developer-oriented data pipeline and workflow orchestration (e.g., Airflow (preferred), dbt, dagster or similar)
  • Knowledge of modern data design and storage patterns (e.g., incremental updating, partitioning and segmentation, rebuilds and backfills)
  • Experience working in Databricks (including delta live tables, data lakehouse patterns, etc.)
  • Experience with cloud computing services (AWS (preferred), GCP, Azure or similar)
  • Experience with data warehousing (e.g., Databricks, Snowflake, Redshift, BigQuery, or similar)
  • Understanding of modern data storage formats and tools (e.g., parquet, ORC, Avro, Delta Lake)

Professional Requirements

  • Must thrive in a fast paced environment and be able to work independently
  • Can work effectively remotely (able to be proactive about managing blockers, proactive on reaching out and asking questions, and participating in team activities)
  • Strong written communication skills on Slack/Chat and in documents
  • You are experienced in writing data design docs (pipeline design, dataflow, schema design)
  • You can scope and breakdown projects, communicate and collaborate progress and blockers effectively with your manager, team, and stakeholders

Nice To Haves:

  • Degree in a quantitative discipline such as computer science, mathematics, statistics, or engineering
  • Experience working with entity data (entity resolution / record linkage)
  • Experience working with data acquisition / data integration
  • Expertise with Python and the Python data stack (e.g., numpy, pandas)
  • Experience with streaming platforms (e.g., Kafka)
  • Experience evaluating data quality and maintaining consistently high data standards across new feature releases (e.g., consistency, accuracy, validity, completeness)

Our Benefits

  • Stock
  • Competitive Salaries
  • Unlimited paid time off
  • Medical, dental, & vision insurance 
  • Health, fitness, and office stipends
  • The permanent ability to work wherever and however you want

Salary: $190K - $220K

No C2C, 1099, or Contract-to-Hire. Recruiters need not apply.

People Data Labs does not discriminate on the basis of race, sex, color, religion, age, national origin, marital status, disability, veteran status, genetic information, sexual orientation, gender identity or any other reason prohibited by law in provision of employment opportunities and benefits.

Qualified Applicants with arrest or conviction records will be considered for Employment in accordance with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.

Personal Privacy Policy for California Residents
https://privacy.peopledatalabs.com/policies?name=personnel-privacy-policy

Top Skills

Java
Python
Scala
The Company
New York, New York
130 Employees
On-site Workplace
Year Founded: 2015

What We Do

At People Data Labs, we know that every company is rapidly transforming to solve problems in a more data and people-driven manner. In order to succeed, companies are forming data departments, pivoting their focus into data acquisition, and doubling down on legal, security, and compliance to protect themselves. For these organizations, clean, rich, and compliant person data is critical and People Data Labs is here to meet that demand.

Today, the People Data Labs platform seeks to enable all companies to build compliant people data solutions. Our sole focus is on building the best data available by integrating thousands of compliantly-sourced datasets into a single, developer-friendly source of truth. Over 2 billion profiles are used by leading companies to enrich recruiting platforms, power AI models, create custom audiences, and more.

Similar Jobs

Arcadia Logo Arcadia

Senior Data Engineer

Big Data • Healthtech • Software • Analytics
Remote
USA
370 Employees

BAE Systems, Inc. Logo BAE Systems, Inc.

Senior Data Engineer [REMOTE

Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Remote
Hybrid
Fort Walton Beach, FL, USA
40000 Employees
76K-128K Annually

Particle Health Logo Particle Health

Healthcare Data Architect / Sr. Data Engineer

Big Data • Healthtech • Information Technology • Software • Analytics • Infrastructure as a Service (IaaS) • Big Data Analytics
Easy Apply
Remote
US
47 Employees

Atlassian Logo Atlassian

Senior Data Engineer

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Remote
San Francisco, CA, USA
11000 Employees
136K-218K Annually

Similar Companies Hiring

Jobba Trade Technologies, Inc. Thumbnail
Software • Professional Services • Productivity • Information Technology • Cloud
Chicago, IL
45 Employees
RunPod Thumbnail
Software • Infrastructure as a Service (IaaS) • Cloud • Artificial Intelligence
Charlotte, North Carolina
53 Employees
Hedra Thumbnail
Software • News + Entertainment • Marketing Tech • Generative AI • Enterprise Web • Digital Media • Consumer Web
San Francisco, CA
14 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account