🔎 I Investigate Databricks Production Incidents

A successful data pipeline can still hide a serious problem. I learned
this after working through production-style Databricks incidents where a
job could finish successfully while the resulting data was stale,
duplicated, incomplete, unexpectedly expensive, or available to the wrong
audience. That changed the way I think about data engineering. I no longer
begin with the question, “Which command should I use?” I
begin with a more useful question: “What evidence do I have, and
which part of the system is responsible for the behavior I am
seeing?”

This article describes the investigation process I use when reasoning
about Databricks pipelines. My goal is not to provide a collection of
shortcuts or claim that one feature solves every production problem. I want
to understand how ingestion, Structured Streaming, Delta tables, data
quality, orchestration, deployment, and governance interact. When an
incident occurs, I try to isolate those responsibilities before changing
anything.

Why I start with evidence instead of assumptions

When an alert arrives, my first instinct can be misleading. A slow
pipeline makes me think about larger compute. A stale table makes me
suspect the transformation. Duplicate records make me think about
deduplication. Those guesses may eventually be correct, but I do not want
my first guess to become my diagnosis.

I start by writing down the observable symptom. I ask when it began,
which datasets are affected, whether the source data arrived, whether the
previous run succeeded, what changed recently, and whether the problem is
reproducible. I separate facts from hypotheses. “The Gold table has not
changed for two hours” is evidence. “Spark is slow” is a hypothesis until I
have something that supports it.

This simple distinction prevents me from changing several unrelated
parts of the system at once. If I make five changes before retesting, even
a successful recovery may teach me very little about the original cause.

I map the incident to a system responsibility

After I record the symptoms, I decide which responsibility is most
likely involved. I use a practical mental map:

  • Ingestion: Did the expected source data arrive, and
    did my ingestion process discover it?
  • Streaming state: Can the query recover, and how does
    it handle event time, state, and late data?
  • Transformation: Did the business logic produce the
    intended rows and columns?
  • Delta table behavior: Are updates, file layout,
    history, and cleanup behaving as expected?
  • Data quality: Did technically valid processing produce
    invalid business data?
  • Orchestration: Did tasks run in the correct order and
    under the intended conditions?
  • Governance: Do identities have the access they
    need—and only that access?
  • Deployment: Is production actually running the
    configuration and resources I intended to deploy?

I do not assume these categories are isolated. A real incident can
involve several of them. The map simply gives me a disciplined place to
begin.

When new files exist but my table does not change

One investigation I often imagine begins with a simple complaint: “The
table is stale.” Before touching the transformation, I verify whether the
source files exist. If the files never arrived, the Databricks
transformation may not be the failing component at all. If the files are
present but the pipeline does not process them, I move my attention to
ingestion.

For a workload designed around incremental file arrival, I inspect the
mechanism responsible for discovering new files. If I am using Auto Loader,
I review the configured source path, schema-related behavior, checkpoint or
state locations where applicable, and recent run information. I also
compare the time the files arrived with the time the pipeline last made
progress.

The important lesson for me is that a stale destination does not
automatically mean a broken transformation. I first prove that data crossed
the ingestion boundary.

When a stream restarts differently than I expect

Streaming incidents force me to think about state. If a query stops and
restarts, I want to know what information survives that interruption. A
durable checkpoint is therefore something I treat as part of the recovery
design rather than an afterthought.

If a restarted query behaves as though it has lost its previous
progress, I inspect its checkpoint configuration before inventing a
complicated explanation. I also avoid casually deleting checkpoint state
during troubleshooting. Removing state can change the behavior I am trying
to understand and can turn an investigation into a new incident.

I ask what recovery guarantee the workload requires, which state is
persisted, and whether a deployment or configuration change caused the
query to use a different location. That gives me a much stronger basis for
action than simply restarting the job repeatedly.

I distinguish event time from processing time

Late data is another source of subtle problems. An event can occur at
one time and arrive at my pipeline later. If I aggregate by event-time
windows, that difference matters.

When state grows continuously in a windowed streaming operation, I
investigate how the query treats late events and whether an appropriate
event-time watermark is part of the design. I do not treat a watermark as a
generic performance switch. I connect it to the business question: how late
can useful data arrive, and how long does the system need to maintain
relevant state?

That means the technical configuration should follow the data contract.
A business process that routinely delivers events late requires a different
discussion from one where records normally arrive within seconds.

When retries create duplicate business effects

Retries are normal in production systems, so I do not design as if every
execution will happen exactly once without interruption. If a workload
partially completes and then runs again, some input may be encountered more
than once.

If I discover duplicate business effects after a retry, I investigate
idempotency. I look for stable identifiers, the write pattern, and the
boundary at which the system decides whether an incoming record represents
something new or something already applied.

This is especially important because a retry mechanism can be
operationally correct while the data-processing logic is not replay-safe.
“The job recovered” and “the data remained correct after recovery” are two
separate statements, and I want evidence for both.

I investigate updates differently from new events

Not every incoming record represents a new fact that should simply be
appended. Some feeds describe changes to existing entities. A customer can
change an address. An order can change status. A source can communicate
inserts, updates, and deletes.

When I see multiple versions of the same business entity in a target
that is supposed to represent current state, I inspect the business key and
the write strategy. A key-based MERGE pattern may be appropriate when I
need different behavior for matching and non-matching records.

For change data capture, I go further. A stable key alone may not tell
me which of several changes should win. I inspect sequencing or ordering
information and determine how deletes are represented. If an older change
is applied after a newer one, the final table can be wrong even though
every individual write succeeded.

Schema changes are evidence, not surprises to ignore

Source systems evolve. A new field can appear. A field can change type.
A producer can begin sending data that does not match the assumptions made
when I first built the pipeline.

If ingestion begins failing immediately after an upstream release, I
compare the old and new source structures. When I find an approved new
column, I investigate the pipeline’s schema-management and schema-evolution
strategy rather than deleting the destination and starting again.

I also distinguish expected evolution from unexpected corruption.
Automatically accepting every possible structural change is not the same
thing as having a controlled schema-evolution policy. I want to know which
changes are acceptable, how they are recorded, and what should happen when
an incompatible change appears.

Successful execution does not prove data quality

One of the most useful investigation habits I have developed is refusing
to equate a green job status with correct business data.

A Spark job can complete without an exception while producing rows that
violate business requirements. A required customer identifier can be null.
An amount can be outside an accepted range. A supposedly unique business
key can appear multiple times.

When users report bad data from a successful run, I inspect explicit
data-quality expectations and how invalid records are handled. Depending on
the requirement, I may need to reject records, quarantine them, record
quality metrics, or stop publication. The exact response depends on the
data contract, but the principle remains the same: technical execution and
business validity are different dimensions of correctness.

I use data layers to narrow the investigation

A layered architecture helps me reason about where a problem entered the
system. If I use Bronze for raw ingestion, Silver for cleaned or conformed
data, and Gold for business-facing outputs, I can compare the same logical
record across those boundaries.

If the raw Bronze record is already wrong, rewriting the Gold
aggregation does not repair the source. If Bronze is correct but Silver is
wrong, I focus on refinement logic. If Silver is correct but Gold is stale,
I inspect publication logic and orchestration.

I find this much more useful than assuming a layer is “good” or “bad”
because of its name. The value of the layers during an incident is that
they give me observable boundaries where I can compare states.

When Delta tables become slow, I inspect the physical problem

Performance investigations require the same discipline. If a Delta table
becomes slower after frequent small writes, I inspect its physical
organization. A large number of small files can create a different problem
from obsolete historical files consuming storage.

I associate OPTIMIZE with improving physical file organization and
compaction. I associate VACUUM with eligible unreferenced-file cleanup
under the relevant retention behavior. I do not treat those operations as
interchangeable merely because both are maintenance operations.

Before cleanup, I also ask whether my organization relies on historical
table states for investigation, recovery, or other workflows. Storage
reclamation is useful, but I do not want a cleanup decision to silently
contradict an operational recovery requirement.

Table history helps me reconstruct what happened

If a production table changes unexpectedly, I want a timeline. Delta
table history can give me evidence about previous table operations. When an
appropriate historical state remains available, comparing versions can help
me determine what changed.

I use that information as evidence rather than as a substitute for
reasoning. A table operation can tell me that something happened, but I may
still need job logs, deployment information, or application context to
understand why it happened.

This is a recurring theme in my investigations: one source of evidence
rarely explains an entire production incident.

When downstream processing is expensive, I ask what actually
changed

A full refresh can be easy to understand, but repeatedly processing
unchanged data can become expensive as datasets grow. If a downstream
workload only needs records affected since its previous successful run, I
investigate whether an incremental pattern fits the requirement.

For suitable Delta workflows, Change Data Feed can be relevant when
row-level changes need to be consumed incrementally and the table is
configured appropriately. I still need to think about the downstream
processing contract, recovery point, and what happens when processing is
retried.

I do not choose incremental processing only because it sounds more
sophisticated. I choose it when the workload requirements justify the
additional state and recovery considerations.

I investigate driver failures by looking for lost distribution

Spark gives me distributed processing, but my own code can accidentally
defeat that advantage. If a job handling hundreds of millions of rows
repeatedly runs out of driver memory, I inspect whether the code collects
large datasets to the driver.

If I find a full DataFrame being collected before a normal Python loop,
that is strong evidence of a scalability problem. When possible, I keep
large transformations expressed through distributed Spark operations rather
than moving the entire dataset onto one machine.

I still examine the wider execution plan when necessary, but I prefer
evidence from the actual data movement over the vague conclusion that “the
cluster is too small.”

Orchestration failures can look like data failures

A stale Gold table may tempt me to rewrite its SQL. But if the Silver
table is current and correct, I ask whether the Gold task actually ran at
the right time.

I inspect workflow dependencies rather than relying only on clock
schedules. If publishing must happen after validation, I want that
relationship represented explicitly. Two independent schedules can appear
to work for weeks and then fail when an upstream task takes longer than
usual.

This is why I treat orchestration as part of data correctness. A correct
transformation executed in the wrong order can still produce an incorrect
or stale result.

Governance incidents require a different kind of evidence

Not every production problem is about rows and files. Sometimes a user
cannot read a table they should be able to access. In another case, a
reporting identity may have permission to modify production data even
though it only needs read access.

When the symptom is authorization, I inspect the relevant Unity Catalog
objects and privileges. I think in terms of least privilege: an identity
should receive the permissions required for its workload without
automatically receiving broader capabilities.

I do not try to solve an authorization problem by changing compute size,
table layout, or streaming configuration. Those changes belong to different
responsibilities.

Deployment differences can explain “works in development”

One of the most frustrating incident reports is “It works in
development.” That statement is useful only if I know whether development
and production are actually equivalent in the ways that matter.

I compare environment-specific parameters, resource definitions,
permissions, runtime configuration, and deployed versions. I prefer
deployment definitions that can be version controlled and reviewed because
they give me evidence of what was intended to change.

Databricks Asset Bundles can support deployment-as-code workflows for
Databricks resources. The broader lesson I take from this is that
reproducible deployment makes incident investigation easier: I can compare
definitions instead of reconstructing production configuration from
memory.

My root-cause workflow

For a serious incident, I use a repeatable sequence. First, I describe
the symptom without embedding a cause in the description. Second, I build a
timeline. Third, I identify the affected system responsibilities. Fourth, I
collect evidence from the relevant runs, tables, configurations,
dependencies, and permissions. Fifth, I create competing hypotheses. Sixth,
I test the cheapest and safest discriminating evidence before making
destructive changes.

After I identify a likely cause, I separate immediate recovery from
long-term prevention. Restoring a pipeline is not always the same task as
preventing recurrence. I may need to correct the current data, repair a
dependency, change an idempotency boundary, strengthen a quality rule, or
improve deployment controls.

Finally, I verify the system end to end. I do not stop merely because
the failed task turns green. I confirm that expected source data was
processed, target data is correct, downstream outputs are current, and
access behavior remains appropriate.

Questions I ask before I close an incident

  • Did I identify the root cause from evidence, or only remove the visible
    symptom?
  • Did the recovery create duplicates, omissions, or unexpected historical
    changes?
  • Are downstream tables and reports current again?
  • Did I preserve the required governance and access boundaries?
  • Could the same failure happen again after the next retry or
    deployment?
  • What monitoring or validation would expose this problem earlier next
    time?

What I want to improve as a data engineer

The most valuable skill I am practicing is not memorizing isolated
Databricks features. I am practicing how to reason under uncertainty. I
want to know which evidence matters, which subsystem owns a behavior, and
which change is justified by the facts I have collected.

That mindset also changes how I design pipelines before an incident
happens. I think about recovery before the first failure, idempotency
before the first retry, data quality before the first bad record,
governance before broad access is granted, and observability before someone
asks why yesterday’s numbers changed.

The investigation scenarios that follow are designed around that idea.
Each case gives me an incident, a piece of evidence, and a decision. I have
to determine what the evidence supports rather than simply recognize a
product name. As the cases become more complex, several subsystems may be
involved at once, just as they are in real production environments.

Note: This is an independent educational article based on
production-style data-engineering scenarios. Product names are used
descriptively. It is not an official certification guide and does not
reproduce certification exam questions.

Leave a Reply

Your email address will not be published. Required fields are marked *

We use cookies and similar technologies to enhance your experience on wobizdu.com, analyze site traffic, personalize content, and deliver relevant ads. Some cookies are essential for the site to function, while others help us improve performance and user experience. You may accept all cookies, decline optional ones, or customize your settings. Review our Privacy Policy to learn more.