Skip to main content
Open-source engineering deep dive

Running Spark against ADLS Gen2 locally with Floci AZ

We implemented the ADLS Gen2 DFS semantics required by hadoop-azure:x.x.x and validated the result through the complete Titan Spark pipeline over abfss://.

Spark 3.5 / 4.x hadoop-azure 3.3 / 3.4 ADLS Gen2 DFS Titan E2E
local compatibility path
real clients
01

Titan / Spark

abfss://...

02

Hadoop ABFS

hadoop-azure:x.x.x

03

ADLS Gen2 DFS semantics

Errors · ETags · list · rename · metadata

04

Floci AZ

Local Azure-compatible storage boundary

Validated

Databricks Runtime compatibility baselines

Validated across three Spark and Hadoop generations

We did not validate the DFS implementation against a single runtime combination. The same Floci AZ DFS path was exercised across our Spark 3.5 and Spark 4.x compatibility baselines.

Baseline 1

DBR 16.4 LTS compatibility

Spark
3.5.2
Hadoop
3.3.4
Client
hadoop-azure:3.3.4
Delta Lake
3.3.1

Baseline 2

DBR 17.3 LTS compatibility

Spark
4.0.0
Hadoop
3.4.1
Client
hadoop-azure:3.4.1
Delta Lake
4.0.0

Baseline 3

DBR 18 LTS compatibility

Spark
4.1.0
Hadoop
3.4.2
Client
hadoop-azure:3.4.2
Delta Lake
4.2.0

The Databricks Runtime references are approximate environment analogues for orientation. They are not claims of bit-for-bit runtime equivalence.

Local emulator choice

Why Floci AZ and not LocalStack or Azurite

Titan depends on cloud-facing contracts, not just storage bytes. In local development we want our normal SDKs, Spark filesystem code and infrastructure tooling to keep running against Azure-shaped endpoints without provisioning Azure infrastructure for every iteration.

Local emulator decision

How the three options compare

The important question was not which emulator has the most features, but which one best preserves Titan's Azure-facing development path.

LocalStack

Not a fit
Cloud boundary
AWS-focused
ADLS Gen2 DFS
× Not the Azure DFS contract
Platform scope
Broad service emulation, but for the wrong cloud boundary
Fit for Titan
No

Azurite

Partial fit
Cloud boundary
Azure Storage
ADLS Gen2 DFS
× Outside documented scope
Platform scope
Blob, Queue and Table storage emulation
Fit for Titan
Partial

Floci AZ

Selected
Cloud boundary
Azure-facing services
ADLS Gen2 DFS
✓ Supported through our DFS contribution
Platform scope
Broader Azure emulator family with real SDK and CLI clients
Fit for Titan
✓ Best fit
✓

Decision: Floci AZ lets us keep the production-facing Spark, Hadoop and Azure client path in the local development loop, while still giving us an implementation we can extend when the required protocol surface is missing.

The first failure

The first Floci AZ / Hadoop ABFS compatibility failure

Hadoop called FileSystem.delete() through the DFS endpoint. Before the patch, that DELETE still went through floci-az's generic Blob delete path. A missing path therefore looked like a missing blob.

Observed before the patch

Blob-style missing object

verified response shape
HTTP 404
Content-Type: application/xml
x-ms-error-code: BlobNotFound

<...>
  <Code>BlobNotFound</Code>
  ...
</...>

Only the fields verified by the regression tooling are shown. We do not claim an exact original XML envelope that was never captured as a fixture.

What Hadoop ABFS needed

Data Lake missing path

DFS response
HTTP 404
Content-Type: application/json
x-ms-error-code: PathNotFound

{
  "error": {
    "code": "PathNotFound",
    "message": "The specified path does not exist."
  }
}

With DFS semantics Hadoop maps the missing delete to the normal filesystem result: false.

How one missing-path DELETE is translated

HTTP response → filesystem result

01

HTTP transport

404 says the resource does not exist.

02

ADLS Gen2 DFS

PathNotFound + JSON identifies the Data Lake service contract.

03

Hadoop ABFS

The client maps the Azure response into Hadoop filesystem semantics.

04

Caller result

fs.delete(path, true) == false

Protocol semantics

Hadoop cares about the contract behind the status code

ABFS interprets Azure-specific error codes, headers and filesystem state across multiple requests. A correct HTTP status with the wrong service semantics can still produce the wrong Hadoop filesystem result. The same contract governs ETags, path types, rename behavior and metadata updates, which is why the implementation had to extend beyond DELETE.

Local parity boundary

Run the real Spark and Hadoop ABFS path locally

Titan still reaches storage through Spark and Hadoop ABFS. Local development changes the endpoint that receives the DFS request, not the client path that creates it.

Spark and Hadoop ABFS path with production and local DFS endpoints Titan calls Spark, Spark uses Hadoop ABFS and Hadoop emits the ADLS Gen2 DFS request. The same client path can terminate at Azure Data Lake Storage in production or at a loopback Docker route into floci-az locally. UNCHANGED CLIENT PATH APPLICATION Titan Pipeline + file formats EXECUTION Apache Spark 3.5.2 / 4.0.0 / 4.1.0 FILESYSTEM CLIENT Hadoop ABFS hadoop-azure:x.x.x ADLS GEN2 DFS AUTHORITY *.dfs.core.windows.net Azure-shaped hostname and request semantics PRODUCTION ENDPOINT Azure Data Lake Storage Gen2 <account>.dfs.core.windows.net Real Azure DFS service LOCAL ENDPOINT Loopback / Docker route → floci-az :4577 devstoreaccount1.dfs.core.windows.net Local DFS-compatible service

The emulator strengthens the inner development loop. It does not replace integration testing against Azure.

How we achieved it

Map the ABFS contract instead of chasing one failure at a time

After a few cycles of fixing one operation only to expose the next one, we were fed up with solving one small piece of the puzzle at a time.

We took hadoop-azure:x.x.x and hadoop-common:x.x.x:

Traced the ABFS code paths

Reverse engineered the DFS calls

Extended Floci AZ DFS handling

That changed the work from reactive debugging into a bounded compatibility pass. We knew which filesystem operations were reachable, which service semantics Hadoop parsed and which states had to survive across calls.

ABFS compatibility workflow

From client analysis to Titan E2E

01

Trace the client

Follow Hadoop filesystem methods into the concrete ABFS request paths.

02

Map the reachable DFS surface

Identify filesystem lifecycle, path I/O, listing, metadata, ACL, leases, rename and delete.

03

Match client-observed semantics

Implement error envelopes, ETags, resource types, continuation and state transitions.

04

Add regression tests

Turn each discovered contract into repository-level coverage.

05

Run the real Hadoop client

Use AzureBlobFileSystem as the client-level compatibility check.

06

Run Titan end to end

Validate every supported ABFSS workflow through Source → Semantic.

Implemented DFS surface

The DFS contract required by hadoop-azure and hadoop-common

The result is broader than the original DELETE issue because the real ABFS client depends on a coherent DFS filesystem contract. We implemented the operations exercised by the combined hadoop-azure and hadoop-common client path, then verified that contract across the Spark 3.5.2, 4.0.0 and 4.1.0 compatibility baselines.

Filesystem

Create and properties

Filesystem create/delete plus get/set filesystem properties and root status.

Path I/O

Create read append flush

File and directory creation, Path Read, sequential append and flush.

Conditions

ETags and overwrite

Conditional create and Hadoop's 409 → status/ETag → If-Match overwrite flow.

Namespace

List status rename delete

Exact-file listing, continuation handling, recursive delete and file/directory rename.

Metadata

Properties and XAttrs

Metadata updates round-trip without changing the underlying file bytes.

Access control

ACL owner permission

Wire-compatible owner, group, permission and ACL metadata plus checkAccess.

HNS + leases

Autodetection and Path Lease

HNS detection, root ACL/status and acquire/renew/release/break/change lease flows.

Error contract

DFS JSON normalization

Data Lake error codes and JSON envelopes are normalized before Hadoop receives them.

Implementation choices

Separate DFS requests from Blob behavior

The important architectural change is service separation. A request for *.dfs.core.windows.net has to be identified before storage behavior is selected. Unsupported DFS updates are rejected explicitly rather than falling through to generic Blob mutation logic.

Fail closed

Metadata calls must never become file uploads

Hadoop metadata operations often have empty request bodies. If an unknown DFS PUT fell through to generic PutBlob, the emulator could replace an existing file with zero bytes.

Rule: supported DFS operation or explicit DFS error. No silent Blob fallback.

HTTP/2 routing

The DFS hostname can arrive through HTTP/2 authority

The routing filter captures the resolved request URI host before async dispatch. That works for HTTP/2 authority as well as HTTP/1.1, with Host as a fallback.

Capture the DFS host before async dispatch

This is the actual branch code. The resolved host is captured while the JAX-RS request context is available, persisted on AzureRequest and then normalized by BlobServiceHandler.

actual implementation excerpt
// AzureRoutingFilter.java
String h = requestContext.getUriInfo().getRequestUri().getHost();
if (h == null || h.isBlank()) {
    h = requestContext.getHeaders().getFirst("Host");
    if (h == null) {
        h = requestContext.getHeaders().getFirst("host");
    }
}
final String capturedHost = h;

// BlobServiceHandler.java
private static boolean isDataLakeRequest(AzureRequest request) {
    String host = request.host();
    if (host == null || host.isBlank()) {
        return false;
    }
    String normalizedHost = host.trim().toLowerCase(Locale.ROOT);
    int portSeparator = normalizedHost.indexOf(':');
    if (portSeparator >= 0) {
        normalizedHost = normalizedHost.substring(0, portSeparator);
    }
    return normalizedHost.endsWith(".dfs.core.windows.net");
}

Inside floci-az

How a DFS request moves through the implementation

The request is classified as Data Lake traffic, dispatched to the DFS path implementation, applied to local state and translated back into the response contract Hadoop expects.

Hadoop ABFS request flow through the floci-az DFS implementation Hadoop ABFS sends an ADLS Gen2 DFS request. AzureRoutingFilter captures the host, BlobServiceHandler classifies the DFS operation, DataLakePathOperations applies path semantics, storage state is updated and AzureErrorResponse creates the correct DFS response. Unsupported DFS updates fail closed. CLIENT FLOCI-AZ REQUEST PIPELINE STATE + RESPONSE 1 · HADOOP ABFS DFS request create · status · list · rename · delete 2 · ROUTING AzureRoutingFilter Capture request URI host / Host fallback Persist authority on AzureRequest 3 · DISPATCH BlobServiceHandler method + action + query + headers select filesystem or path operation 4 · DFS SEMANTICS DataLakePathOperations create · read · append · flush list · properties · ACL · access rename · delete · continuation 5 · STATE StorageBackend + leases bytes · ETags · resource type · metadata Path Lease state where applicable 6 · DFS RESPONSE Azure-compatible result status · headers · JSON error semantics FAIL CLOSED Unsupported DFS update Return DFS error. Never PutBlob.

Repository compatibility test

Use the real Hadoop ABFS client

REST tests prove endpoint behavior. This test proves that the endpoint semantics compose into a filesystem the actual Hadoop 3.3.4 client can use.

HNS + root status
Conditional overwrite
XAttrs + exact listing
Rename + missing delete

Actual test harness

Preserve Hadoop's DFS hostname and wire shape

The Docker compatibility setup maps devstoreaccount1.dfs.core.windows.net to loopback. A transparent TCP relay forwards port 80 to floci-az:4577, so Hadoop still emits its normal DFS request.

Hadoop 3.3.4 defaults fs.azure.always.use.https=true, so the test disables HTTPS only for this plaintext loopback harness.

loopback-only forwarder
void start() throws IOException {
    serverSocket = new ServerSocket(
            listenPort,
            50,
            InetAddress.getByName("127.0.0.1"));
    running.set(true);
    executor.submit(this::acceptLoop);
}

Binding explicitly to 127.0.0.1 keeps the test relay off the container's other interfaces.

HadoopAbfsCompatibilityTest.java · actual excerpt
String filesystem =
        "hadoop-" + UUID.randomUUID().toString().replace("-", "").substring(0, 12);
URI uri = URI.create("abfs://" + filesystem + "@" + ACCOUNT_HOST);

assertEquals("3.3.4", VersionInfo.getVersion());

Configuration conf = new Configuration();
conf.set("fs.abfs.impl",
        "org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem");
conf.set("fs.azure.account.key." + ACCOUNT_HOST, EmulatorConfig.DEV_KEY);
conf.setBoolean("fs.azure.always.use.https", false);
conf.setBoolean("fs.azure.createRemoteFileSystemDuringInitialization", true);
conf.setBoolean("fs.azure.enable.conditional.create.overwrite", true);
conf.unset("fs.azure.account.hns.enabled." + ACCOUNT_HOST);
conf.unset("fs.azure.account.hns.enabled");

try (FileSystem fs = FileSystem.newInstance(uri, conf)) {
    assertTrue(fs.getFileStatus(new Path("/")).isDirectory());

    Path dir = new Path("/compat");
    Path file = new Path(dir, "file.txt");
    assertTrue(fs.mkdirs(dir));

    write(fs, file, "first");
    write(fs, file, "second");
    assertEquals("second", read(fs, file));

    byte[] xattr = "v8".getBytes(StandardCharsets.UTF_8);
    fs.setXAttr(file, "user.floci_smoke", xattr);
    assertArrayEquals(xattr, fs.getXAttr(file, "user.floci_smoke"));

    FileStatus[] exact = fs.listStatus(file);
    assertEquals(1, exact.length);

    Path renamed = new Path(dir, "renamed.txt");
    assertTrue(fs.rename(file, renamed));

    assertFalse(fs.delete(new Path(dir, "missing.txt"), true));
    assertTrue(fs.delete(dir, true));
}

Application-level acceptance test

Validate the complete Titan pipeline on ABFSS

The full Titan application-level E2E gate shown here uses Spark 3.5.2 + Hadoop 3.3.4 + SecureAzureBlobFileSystem over abfss://. The same DFS implementation was also validated against the Spark 4.0.0 / Hadoop 3.4.1 and Spark 4.1.0 / Hadoop 3.4.2 compatibility baselines listed above.

Titan repeatedly crosses the filesystem boundary while discovering source files, reading multiple formats, writing intermediate datasets, listing results and committing Spark output. The E2E run verifies that those contracts work together across the complete pipeline.

Source formats tested end to end

Nine input workflows. One ABFSS-backed Titan pipeline.

Avro CSV CSV.GZ Delimited JSON OCR Parquet XML XLSX
Complete Titan ABFSS end-to-end data flow All nine tested source formats enter the Source layer and continue through Bronze, Silver, Technical Decoupling, Gold and Semantic. 01 Source 02 Bronze 03 Silver 04 Technical Decoupling 05 Gold 06 Semantic
930 / 930floci-az repository tests
263 tests · 0 failuresJava compatibility suite including real Hadoop ABFS 3.3.4
9 / 9 workflowsTitan ABFSS E2E through Source → Semantic

Reproduce the validation

Keep the compatibility claim executable

The useful outcome is not a one-off green run. The repository carries the Hadoop compatibility path so a future change can be checked against the same client contract.

repository suite
./mvnw test

# final result
Tests run: 930
Failures: 0
Errors: 0
Skipped: 0
java compatibility suite
make test-java-compat

# final result
Tests run: 263
Failures: 0
Errors: 0
Skipped: 2
BUILD SUCCESS

Engineering takeaway

Move filesystem failures into the local feedback loop

The main benefit is not that Azure suddenly runs on a laptop. The benefit is that the same Spark and Hadoop filesystem path can now be exercised before a deployment reaches Azure, which moves protocol and integration failures much closer to the code change that caused them.

Real client behavior locally. Spark and Hadoop ABFS stay in the execution path instead of being replaced by storage mocks.

Earlier protocol feedback. Wrong DFS errors, listing behavior, metadata updates and rename semantics can fail in the local or CI loop.

Cleaner Azure validation. Cloud tests can focus on identity, private networking, service limits and production-specific behavior.

The result is a shorter feedback cycle without pretending that a local emulator is a substitute for Azure itself.

Shorter feedback loop

Catch the storage contract before cloud deployment

Before
change code → build → deploy Azure → run pipeline → discover ABFS issue → debug remotely
Now
change code → run local Titan + Hadoop → catch contract issue

FAQ

Local ADLS Gen2, Hadoop ABFS and Spark questions

Practical answers about the protocol choices, test harness and scope of the Floci AZ contribution.

Can Apache Spark use ADLS Gen2 locally without a real Azure Storage Account?

Yes, provided the emulator reproduces the DFS wire semantics that Hadoop ABFS expects. The important part is not only accepting the same URLs; service-specific errors, ETags, resource types, listing behaviour and path update semantics must also match the client contract.

Why did BlobNotFound versus PathNotFound break FileSystem.delete()?

Hadoop ABFS interprets the ADLS service error code and JSON body. A Blob-style 404 is not semantically equivalent to a DFS PathNotFound response. With the correct DFS response, deleting a missing path maps to the normal Hadoop filesystem result: false.

Which Hadoop implementations were validated?

The repository compatibility suite directly uses hadoop-common 3.3.4 and hadoop-azure 3.3.4 with the real org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem class. Additional compatibility baselines cover Hadoop 3.4.1 with hadoop-azure 3.4.1 and Hadoop 3.4.2 with hadoop-azure 3.4.2 alongside Spark 4.0.0 and 4.1.0.

Was abfss:// tested too?

Yes. The Docker compatibility harness uses abfs:// over loopback HTTP so it can preserve the DFS host cleanly. External Spark validation used org.apache.hadoop.fs.azurebfs.SecureAzureBlobFileSystem and abfss:// paths.

Why is fs.azure.always.use.https=false set in the compatibility test?

In the Hadoop 3.3.4 repository compatibility harness, this setting defaults to true even for abfs://. That local harness deliberately forwards plaintext loopback traffic on port 80, so HTTPS is disabled only for that test. It does not alter Floci's runtime routing behaviour.

Does this implement complete ADLS Gen2 POSIX security?

No. Owner, group, permission and ACL metadata round-trip through the DFS/Hadoop wire protocol, but the emulator does not attempt to reproduce the complete Azure POSIX authorization engine.

How was the final implementation validated?

The final branch passed 930 repository tests and 263 Java compatibility tests with zero failures. The repository suite directly exercises the real Hadoop 3.3.4 ABFS client, the Titan application-level E2E gate runs Spark 3.5.2 over abfss:// across nine source workflows, and additional compatibility baselines validate the DFS implementation with Spark 4.0.0 / Hadoop 3.4.1 and Spark 4.1.0 / Hadoop 3.4.2.

Where is the open-source implementation?

The implementation was developed for the Floci AZ project and is available on the public feature branch linked from this article while the upstream contribution is reviewed.

Build closer to production

Make local data-platform tests exercise the real integration boundary

We design Spark, Databricks and Azure data platforms with repeatable infrastructure and test boundaries that cover more than application code alone.

A stronger test boundary

Application and Spark behaviour
Real Hadoop filesystem client
ADLS Gen2 DFS wire semantics
Reproducible local + CI validation