Skip to content
    Free template · Updated 2 October 2026

    Data engineer interview questions: 24 questions and a scorecard for interviewers

    The short answer

    Good data engineer interview questions test how the candidate designs and runs reliable data pipelines, models data for a warehouse or lakehouse, writes efficient SQL, tests data quality, chooses between batch and streaming, and controls cloud cost and access to sensitive data. Ask them to walk through one pipeline they built from the source systems to the people who use the data, and listen for how they handled failures, late or bad data and upstream changes, and which parts were their own work. A recruiter screen before the client’s technical round should confirm the stack the client uses, such as the cloud platform, warehouse, orchestration tool and languages, plus work arrangement, on-call, right to work and start date. Score every candidate against the same five criteria so the decision rests on evidence.

    All 24 questions with why you ask each one, what a strong answer shows and follow-ups, plus the data engineer scorecard and rating guide. Free to download and adapt; no sign-up needed.

    When to use it

    When to use these questions

    Data engineers build the pipelines, data models and platforms that analysts, data scientists and applications depend on, so an interview needs to show how someone designs for reliability, data quality, cost and security, not only which tools appear on their resume. The strongest evidence comes from specific pipelines and incidents: where the data came from, how it was modeled, what broke and what they changed so it wouldn’t break the same way again.

    Stacks vary widely between companies, so agree the must-haves with the hiring manager first: the cloud platform and warehouse, orchestration and transformation tools, batch or streaming work, languages, the data volumes involved and whether the role includes on-call. Strong engineers can usually move between similar tools, so weigh how well the candidate understands the underlying ideas as well as exact tool matches, and leave hands-on coding and design assessment to the client’s engineers.

    24 questions

    24 data engineer interview questions

    Grouped by what they test. Pick the questions that match the role, ask every candidate the same ones in the same order, and score each answer against the scorecard below.

    Pipelines, ETL and ELT

    Start with one pipeline the candidate built or owns, end to end. Separate their design decisions from the platform they inherited.

    1. Question 1: Walk me through a data pipeline you built or own, from the source systems to the people who use the data. Which parts did you design yourself?

      Why ask it:
      Establishes real ownership and how the candidate thinks end to end.
      A strong answer shows:
      Sources, ingestion, transformation, storage and consumers explained clearly, with their own design decisions named.
      Follow-up:
      What breaks most often in that pipeline, and what have you done about it?
    2. Question 2: When do you transform data before loading it, and when do you load it raw and transform it inside the warehouse?

      Why ask it:
      Tests understanding of the trade-offs between ETL and ELT.
      A strong answer shows:
      Reasons tied to warehouse compute, the value of keeping raw data for reprocessing, sensitive fields that should be removed early, cost and the skills of the team.
    3. Question 3: How do you make a pipeline safe to rerun after a failure without creating duplicates or gaps?

      Why ask it:
      Pipelines that can be rerun safely prevent silent data errors.
      A strong answer shows:
      Idempotent loads such as merges or partition overwrites, checkpoints or watermarks, and a backfill process they have actually used.
    4. Question 4: A source team renames one column and drops another without telling you. How should your pipelines react, and how do you stop it happening again?

      Why ask it:
      Unannounced upstream changes are a common cause of broken data.
      A strong answer shows:
      Schema checks that fail loudly or quarantine the data, an alert to the right owner, and an agreement or data contract with the source team about future changes.
    5. Question 5: How do you handle records that arrive late or out of order?

      Why ask it:
      Tests correctness thinking for real-world data.
      A strong answer shows:
      The difference between event time and processing time, reprocessing windows or explicit late-data handling, and updating downstream tables and reports when late records land.

    Data modeling, warehouses and SQL

    Ask about tables other people query every day. Good modeling and SQL make data easy to use correctly and cheap to query.

    1. Question 6: How did you model the data in a warehouse you worked on, and why did you choose that approach?

      Why ask it:
      Modeling choices decide how easy data is to use and trust.
      A strong answer shows:
      A clear approach, such as dimensional models with facts and dimensions or layered raw, cleaned and business-ready tables, chosen for the people querying it and the questions they ask.
    2. Question 7: A customer changes their address or subscription plan, and reports need to show both the old and the new values. How do you model that?

      Why ask it:
      Tests practical understanding of keeping history in a warehouse.
      A strong answer shows:
      Keeping history with effective dates or a slowly changing dimension where it matters, overwriting where it doesn’t, and the effect on joins and reports.
    3. Question 8: Analysts say queries on one of your core tables have become slow and expensive. How do you investigate and fix it?

      Why ask it:
      Tests SQL and warehouse performance skills.
      A strong answer shows:
      Reading the query plan, checking how much data is scanned, partitioning or clustering on common filters, pre-aggregated tables and fixing inefficient joins, with a before-and-after result.
      Follow-up:
      How did you confirm the change helped without breaking anything downstream?
    4. Question 9: Describe a SQL transformation that other teams depend on. How did you structure it so others could read, test and change it safely?

      Why ask it:
      Shared transformations need to be maintainable as well as correct.
      A strong answer shows:
      Modular steps such as CTEs or layered models, clear naming, tests on keys and business rules, version control and code review.
    5. Question 10: When would you choose a data warehouse, a data lake or a lakehouse, and which have you used in practice?

      Why ask it:
      Tests architectural understanding of storage options.
      A strong answer shows:
      Choices tied to data types, query patterns, cost, governance and team skills, with hands-on examples rather than vendor slogans.

    Orchestration, reliability and data quality

    Pipelines fail. Look for engineers who find problems before users do and fix the cause, not only the symptom.

    1. Question 11: How do you schedule and orchestrate pipelines that depend on one another?

      Why ask it:
      Tests reliability across chains of dependent jobs.
      A strong answer shows:
      An orchestration tool such as Airflow, Dagster or a cloud scheduler, dependencies defined explicitly, limited retries, alerting and clear ownership of each job.
    2. Question 12: Which data quality checks do you build into a pipeline, and what happens when one fails?

      Why ask it:
      Quality checks stop bad data before it reaches reports and models.
      A strong answer shows:
      Checks on row counts, nulls, uniqueness, accepted values, freshness and reconciliation with the source, plus a clear rule on whether to block, quarantine or warn.
    3. Question 13: Tell me about a time bad or missing data reached a dashboard or model that people relied on. What did you do, and what changed afterwards?

      Why ask it:
      Tests ownership and learning when data fails.
      A strong answer shows:
      Telling affected users quickly, fixing and backfilling the data, a blameless review and a new check or process that prevents a repeat.
      Follow-up:
      How did the people using that data find out, and how quickly?
    4. Question 14: How do you test pipeline code before it reaches production?

      Why ask it:
      Shows engineering discipline beyond writing queries.
      A strong answer shows:
      Unit tests for transformation logic, runs on sample or staging data, code review and automated checks in continuous integration.
    5. Question 15: What did on-call support for data pipelines look like in your current or last role?

      Why ask it:
      Shows how the candidate supports production pipelines when they fail.
      A strong answer shows:
      Actionable alerts, runbooks, priorities set by business impact, and work to reduce repeat alerts over time.

    Streaming, cloud, security and collaboration

    These questions test judgment about cost, freshness and access, and how well the candidate serves the people who use the data.

    1. Question 16: Tell me about a time you chose between batch and streaming for a use case. What decided it?

      Why ask it:
      Streaming adds cost and complexity that a use case doesn’t always need.
      A strong answer shows:
      How fresh the data truly needed to be, volumes, cost, team skills and failure handling, with honesty about when batch was enough.
    2. Question 17: What have you done to keep the cost of a cloud data platform under control?

      Why ask it:
      Warehouse and compute costs can grow quickly without attention.
      A strong answer shows:
      Monitoring spend by team or job, right-sizing compute, scheduling, reducing data scanned, storage lifecycle rules and a specific saving they made.
    3. Question 18: How do you control who can see sensitive data, such as personal or financial information, in a warehouse?

      Why ask it:
      Data engineers often build and enforce the access rules.
      A strong answer shows:
      Role-based access with least privilege, masking or tokenizing sensitive columns, separate environments, audit logs and agreeing with data owners and security who should see what.
    4. Question 19: How do you work with analysts and data scientists to decide what data to build and how to shape it?

      Why ask it:
      Data engineering succeeds only when the people downstream can use what is built.
      A strong answer shows:
      Asking which questions and models the data must serve, agreeing definitions and freshness, documenting tables and building shared models instead of one-off extracts.
    5. Question 20: Explain a pipeline or data platform you worked on as if I were a business manager with no technical background.

      Why ask it:
      Tests communication directly, and a non-technical recruiter can judge it.
      A strong answer shows:
      Plain language about what data it moves, who relies on it and what would go wrong without it, free of jargon.

    Must-haves and logistics

    Ask these of every candidate before the client’s technical round, and check the answers against the resume.

    1. Question 21: Which cloud platforms, warehouses, orchestration and transformation tools and programming languages have you used in production in the last two years, and at roughly what data volumes?

      Why ask it:
      Confirms hands-on experience with the client’s stack at a similar scale.
      A strong answer shows:
      Specific tools tied to real pipelines and recent dates, volumes they can explain, and honesty about tools used only in courses or side projects.
    2. Question 22: This role is [remote, hybrid or on-site, location, and any on-call rotation]. Does that suit you, and what is your notice period or earliest start date?

      Why ask it:
      Rules out arrangement and timing mismatches before the client invests interview time.
      A strong answer shows:
      A clear yes, or the specific constraint, and a firm start date.
    3. Question 23: Are you legally authorized to work in [country] for this employer, and will you need visa sponsorship now or in the future?

      Why ask it:
      Confirms eligibility; ask every candidate the same question in the same way.
      A strong answer shows:
      A direct answer, recorded the same way for every candidate.
    4. Question 24: What pay range are you looking for in this role?

      Why ask it:
      Checks fit with the client’s budget without asking about pay history.
      A strong answer shows:
      A realistic range you can compare with the client’s budget.
    Example rubric

    Data engineer interview scorecard

    Five criteria for this role, with what a score of 1, 3 and 5 looks like. Scores of 2 and 4 sit between them.

    Data engineer interview scorecard
    CriterionWhat it meansScore 1 looks likeScore 3 looks likeScore 5 looks like
    Pipeline design and reliabilityBuilds pipelines that recover cleanly and handle real-world data.Describes tools but not how data flows through a pipeline they built, or reruns jobs without thinking about duplicates.Explains a pipeline end to end, with safe reruns and sensible orchestration.Designs for failure, late data and upstream change from the start, with backfills and clear ownership.
    Data modeling and SQLModels data so it is easy to use and efficient to query.No clear modeling approach and no example of improving query performance.A sensible warehouse model, history handled where needed and a real performance fix.Chooses models for how the data is used, writes maintainable shared transformations and cuts query cost and time measurably.
    Data quality and testingCatches bad data and broken code before users do.Learns about data problems from users and has no testing habits.Builds standard quality checks and tests transformation code before release.Layers checks by risk, decides when to block or warn, runs blameless reviews and prevents repeat incidents.
    Cloud cost and securityKeeps platforms affordable and sensitive data protected.Has never looked at platform costs and treats access control as someone else’s job.Monitors spend and applies role-based access to sensitive data.Delivers specific savings, applies least privilege and masking, and works with security and data owners on access.
    Collaboration and communicationWorks with analysts and data scientists and explains technical work plainly.Builds what is asked without understanding how the data will be used, or relies on jargon.Agrees definitions and freshness with users and explains work clearly.Shapes data around the questions people need answered, documents it well and is trusted by downstream teams.
    Scoring

    The 1–5 rating scale

    The same scale for every criterion and every candidate.

    The 1–5 rating scale
    ScoreLevelWhat it means
    1Well below requirementNo relevant evidence, or an answer that contradicts the requirement.
    2Below requirementPartial evidence with important gaps.
    3Meets requirementClear, relevant evidence at the level the role needs.
    4Above requirementStrong, specific evidence beyond the expected level.
    5ExceptionalRepeated high-quality evidence with clear impact.
    How to use it

    How to run the interview with these data engineer interview questions

    1. 01

      Step 01

      Agree the must-haves first

      Confirm the essential credentials, experience and availability with the hiring manager or client before any interviews.
    2. 02

      Step 02

      Pick 8 to 12 questions

      Take the must-have questions, then the questions that test what this role needs most. Use the same set, in the same order, for every candidate.
    3. 03

      Step 03

      Ask for real examples

      When you hear “we” or “I would”, ask what the candidate personally did, and what happened in the end.
    4. 04

      Step 04

      Score before you discuss

      Rate each criterion on the scorecard with the evidence behind it, then compare with other interviewers.
    5. 05

      Step 05

      Verify before you submit

      Check licenses, certifications and right to work against the original source before you put the candidate forward.
    Watch out

    Red flags, and questions not to ask

    • Lists many tools but can’t describe how data actually moves through a pipeline they built.
    • Has no data quality checks and usually hears about bad data from the people using it.
    • Reruns failed jobs without considering duplicates, gaps or downstream tables.
    • Treats access to sensitive data as the security team’s problem rather than part of the design.
    • Blames source teams or analysts for every data problem, with no example of improving the relationship.
    • Age, marital or family status, pregnancy or plans for children, religion, ethnicity or national origin, sexual orientation or gender identity. These are protected characteristics under the UK Equality Act 2010 and US federal law, and they say nothing about whether someone can do the job.
    • Health, sickness absence or disability before an offer. You can ask whether the candidate needs any adjustments for the interview, and whether they can do the essential tasks of the job.
    Run it in Beatview

    Screen data engineer applicants before the first call

    Add these questions to a Beatview AI interview and every applicant answers them on video or audio, with the same time limit. Beatview scores each answer against your criteria and shows the reasoning, and you can share the shortlist with your client through a password-protected link. AI interviews are on the Pro plan; the Free plan screens resumes for one active job.

    FAQ

    Data engineer interview questions: frequently asked questions

    Still deciding?

    Bring a live vacancy and we’ll walk through where automation ends and recruiter review begins.

    Ask the candidate to walk through one pipeline from source to consumer, then ask questions that test ETL and ELT design, data modeling and SQL performance, orchestration and reliability, data quality and testing, batch versus streaming, cloud cost, security and how they work with analysts and data scientists. Add must-have questions on the client’s stack, work arrangement, on-call, right to work and start date, and ask every candidate the same questions in the same order.

    Ask the candidate to explain a pipeline in plain language: where the data comes from, who uses it and what happens when it breaks. You can judge clarity, ownership and whether they mention testing, monitoring and data quality without being an engineer. Use the listening notes for the technical questions, confirm which tools they have used in production and when, and leave coding and design exercises to the client’s technical round.

    It depends on the client’s timeline. A candidate who knows the client’s exact cloud platform, warehouse and orchestration tool will be productive sooner, but the underlying skills of pipeline design, modeling, SQL, testing and reliability carry over between similar tools. Agree with the hiring manager which tools are essential and which can be learned, and screen for understanding as well as names on a resume.

    The same criteria for every candidate, such as pipeline design and reliability, data modeling and SQL, data quality and testing, cloud cost and security, and collaboration and communication, each with a 1–5 rating, the evidence behind the rating, must-have checks on stack and logistics, and a clear recommendation.

    Download

    Get the data engineer interview questions template

    All 24 questions with why you ask each one, what a strong answer shows and follow-ups, plus the data engineer scorecard and rating guide.

    Opens in Excel, Google Sheets or Numbers. Version 2 October 2026.

    Download the free template (CSV)
    Start free

    Put the template to work on a live role.

    Beatview screens every application against your criteria and interviews the shortlist with the same structured questions. The Free plan covers one active job; AI interviews are on Pro.

    • Free plan with one active role
    • Runs alongside your ATS
    • Recruiters keep every decision

    Page last reviewed by the Beatview team.