Running Spark against ADLS Gen2 locally with Floci AZ
We implemented the ADLS Gen2 DFS semantics required by hadoop-azure:x.x.x and validated the result through the complete Titan Spark pipeline over abfss://.
Titan / Spark
abfss://...
Hadoop ABFS
hadoop-azure:x.x.x
ADLS Gen2 DFS semantics
Errors · ETags · list · rename · metadata
Floci AZ
Local Azure-compatible storage boundary
Databricks Runtime compatibility baselines
Validated across three Spark and Hadoop generations
We did not validate the DFS implementation against a single runtime combination. The same Floci AZ DFS path was exercised across our Spark 3.5 and Spark 4.x compatibility baselines.
Baseline 1
DBR 16.4 LTS compatibility
- Spark
- 3.5.2
- Hadoop
- 3.3.4
- Client
- hadoop-azure:3.3.4
- Delta Lake
- 3.3.1
Baseline 2
DBR 17.3 LTS compatibility
- Spark
- 4.0.0
- Hadoop
- 3.4.1
- Client
- hadoop-azure:3.4.1
- Delta Lake
- 4.0.0
Baseline 3
DBR 18 LTS compatibility
- Spark
- 4.1.0
- Hadoop
- 3.4.2
- Client
- hadoop-azure:3.4.2
- Delta Lake
- 4.2.0
The Databricks Runtime references are approximate environment analogues for orientation. They are not claims of bit-for-bit runtime equivalence.
Local emulator choice
Why Floci AZ and not LocalStack or Azurite
Titan depends on cloud-facing contracts, not just storage bytes. In local development we want our normal SDKs, Spark filesystem code and infrastructure tooling to keep running against Azure-shaped endpoints without provisioning Azure infrastructure for every iteration.
Local emulator decision
How the three options compare
The important question was not which emulator has the most features, but which one best preserves Titan's Azure-facing development path.
LocalStack
Not a fit- Cloud boundary
- AWS-focused
- ADLS Gen2 DFS
- × Not the Azure DFS contract
- Platform scope
- Broad service emulation, but for the wrong cloud boundary
- Fit for Titan
- No
Azurite
Partial fit- Cloud boundary
- Azure Storage
- ADLS Gen2 DFS
- × Outside documented scope
- Platform scope
- Blob, Queue and Table storage emulation
- Fit for Titan
- Partial
Floci AZ
Selected- Cloud boundary
- Azure-facing services
- ADLS Gen2 DFS
- ✓ Supported through our DFS contribution
- Platform scope
- Broader Azure emulator family with real SDK and CLI clients
- Fit for Titan
- ✓ Best fit
Decision: Floci AZ lets us keep the production-facing Spark, Hadoop and Azure client path in the local development loop, while still giving us an implementation we can extend when the required protocol surface is missing.
The first failure
The first Floci AZ / Hadoop ABFS compatibility failure
Hadoop called FileSystem.delete() through the DFS endpoint. Before the patch, that DELETE still went through floci-az's generic Blob delete path. A missing path therefore looked like a missing blob.
Observed before the patch
Blob-style missing object
HTTP 404
Content-Type: application/xml
x-ms-error-code: BlobNotFound
<...>
<Code>BlobNotFound</Code>
...
</...>
Only the fields verified by the regression tooling are shown. We do not claim an exact original XML envelope that was never captured as a fixture.
What Hadoop ABFS needed
Data Lake missing path
HTTP 404
Content-Type: application/json
x-ms-error-code: PathNotFound
{
"error": {
"code": "PathNotFound",
"message": "The specified path does not exist."
}
}
With DFS semantics Hadoop maps the missing delete to the normal filesystem result: false.
How one missing-path DELETE is translated
HTTP response → filesystem result
HTTP transport
404 says the resource does not exist.
ADLS Gen2 DFS
PathNotFound + JSON identifies the Data Lake service contract.
Hadoop ABFS
The client maps the Azure response into Hadoop filesystem semantics.
Caller result
fs.delete(path, true) == false
Protocol semantics
Hadoop cares about the contract behind the status code
ABFS interprets Azure-specific error codes, headers and filesystem state across multiple requests. A correct HTTP status with the wrong service semantics can still produce the wrong Hadoop filesystem result. The same contract governs ETags, path types, rename behavior and metadata updates, which is why the implementation had to extend beyond DELETE.
Local parity boundary
Run the real Spark and Hadoop ABFS path locally
Titan still reaches storage through Spark and Hadoop ABFS. Local development changes the endpoint that receives the DFS request, not the client path that creates it.
The emulator strengthens the inner development loop. It does not replace integration testing against Azure.
How we achieved it
Map the ABFS contract instead of chasing one failure at a time
After a few cycles of fixing one operation only to expose the next one, we were fed up with solving one small piece of the puzzle at a time.
We took hadoop-azure:x.x.x and hadoop-common:x.x.x:
Traced the ABFS code paths
Reverse engineered the DFS calls
Extended Floci AZ DFS handling
That changed the work from reactive debugging into a bounded compatibility pass. We knew which filesystem operations were reachable, which service semantics Hadoop parsed and which states had to survive across calls.
ABFS compatibility workflow
From client analysis to Titan E2E
Trace the client
Follow Hadoop filesystem methods into the concrete ABFS request paths.
Map the reachable DFS surface
Identify filesystem lifecycle, path I/O, listing, metadata, ACL, leases, rename and delete.
Match client-observed semantics
Implement error envelopes, ETags, resource types, continuation and state transitions.
Add regression tests
Turn each discovered contract into repository-level coverage.
Run the real Hadoop client
Use AzureBlobFileSystem as the client-level compatibility check.
Run Titan end to end
Validate every supported ABFSS workflow through Source → Semantic.
Implemented DFS surface
The DFS contract required by hadoop-azure and hadoop-common
The result is broader than the original DELETE issue because the real ABFS client depends on a coherent DFS filesystem contract. We implemented the operations exercised by the combined hadoop-azure and hadoop-common client path, then verified that contract across the Spark 3.5.2, 4.0.0 and 4.1.0 compatibility baselines.
Create and properties
Filesystem create/delete plus get/set filesystem properties and root status.
Create read append flush
File and directory creation, Path Read, sequential append and flush.
ETags and overwrite
Conditional create and Hadoop's 409 → status/ETag → If-Match overwrite flow.
List status rename delete
Exact-file listing, continuation handling, recursive delete and file/directory rename.
Properties and XAttrs
Metadata updates round-trip without changing the underlying file bytes.
ACL owner permission
Wire-compatible owner, group, permission and ACL metadata plus checkAccess.
Autodetection and Path Lease
HNS detection, root ACL/status and acquire/renew/release/break/change lease flows.
DFS JSON normalization
Data Lake error codes and JSON envelopes are normalized before Hadoop receives them.
Implementation choices
Separate DFS requests from Blob behavior
The important architectural change is service separation. A request for *.dfs.core.windows.net has to be identified before storage behavior is selected. Unsupported DFS updates are rejected explicitly rather than falling through to generic Blob mutation logic.
Fail closed
Metadata calls must never become file uploads
Hadoop metadata operations often have empty request bodies. If an unknown DFS PUT fell through to generic PutBlob, the emulator could replace an existing file with zero bytes.
Rule: supported DFS operation or explicit DFS error. No silent Blob fallback.
HTTP/2 routing
The DFS hostname can arrive through HTTP/2 authority
The routing filter captures the resolved request URI host before async dispatch. That works for HTTP/2 authority as well as HTTP/1.1, with Host as a fallback.
Capture the DFS host before async dispatch
This is the actual branch code. The resolved host is captured while the JAX-RS request context is available, persisted on AzureRequest and then normalized by BlobServiceHandler.
// AzureRoutingFilter.java
String h = requestContext.getUriInfo().getRequestUri().getHost();
if (h == null || h.isBlank()) {
h = requestContext.getHeaders().getFirst("Host");
if (h == null) {
h = requestContext.getHeaders().getFirst("host");
}
}
final String capturedHost = h;
// BlobServiceHandler.java
private static boolean isDataLakeRequest(AzureRequest request) {
String host = request.host();
if (host == null || host.isBlank()) {
return false;
}
String normalizedHost = host.trim().toLowerCase(Locale.ROOT);
int portSeparator = normalizedHost.indexOf(':');
if (portSeparator >= 0) {
normalizedHost = normalizedHost.substring(0, portSeparator);
}
return normalizedHost.endsWith(".dfs.core.windows.net");
}
Inside floci-az
How a DFS request moves through the implementation
The request is classified as Data Lake traffic, dispatched to the DFS path implementation, applied to local state and translated back into the response contract Hadoop expects.
Repository compatibility test
Use the real Hadoop ABFS client
REST tests prove endpoint behavior. This test proves that the endpoint semantics compose into a filesystem the actual Hadoop 3.3.4 client can use.
Actual test harness
Preserve Hadoop's DFS hostname and wire shape
The Docker compatibility setup maps devstoreaccount1.dfs.core.windows.net to loopback. A transparent TCP relay forwards port 80 to floci-az:4577, so Hadoop still emits its normal DFS request.
Hadoop 3.3.4 defaults fs.azure.always.use.https=true, so the test disables HTTPS only for this plaintext loopback harness.
void start() throws IOException {
serverSocket = new ServerSocket(
listenPort,
50,
InetAddress.getByName("127.0.0.1"));
running.set(true);
executor.submit(this::acceptLoop);
}
Binding explicitly to 127.0.0.1 keeps the test relay off the container's other interfaces.
String filesystem =
"hadoop-" + UUID.randomUUID().toString().replace("-", "").substring(0, 12);
URI uri = URI.create("abfs://" + filesystem + "@" + ACCOUNT_HOST);
assertEquals("3.3.4", VersionInfo.getVersion());
Configuration conf = new Configuration();
conf.set("fs.abfs.impl",
"org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem");
conf.set("fs.azure.account.key." + ACCOUNT_HOST, EmulatorConfig.DEV_KEY);
conf.setBoolean("fs.azure.always.use.https", false);
conf.setBoolean("fs.azure.createRemoteFileSystemDuringInitialization", true);
conf.setBoolean("fs.azure.enable.conditional.create.overwrite", true);
conf.unset("fs.azure.account.hns.enabled." + ACCOUNT_HOST);
conf.unset("fs.azure.account.hns.enabled");
try (FileSystem fs = FileSystem.newInstance(uri, conf)) {
assertTrue(fs.getFileStatus(new Path("/")).isDirectory());
Path dir = new Path("/compat");
Path file = new Path(dir, "file.txt");
assertTrue(fs.mkdirs(dir));
write(fs, file, "first");
write(fs, file, "second");
assertEquals("second", read(fs, file));
byte[] xattr = "v8".getBytes(StandardCharsets.UTF_8);
fs.setXAttr(file, "user.floci_smoke", xattr);
assertArrayEquals(xattr, fs.getXAttr(file, "user.floci_smoke"));
FileStatus[] exact = fs.listStatus(file);
assertEquals(1, exact.length);
Path renamed = new Path(dir, "renamed.txt");
assertTrue(fs.rename(file, renamed));
assertFalse(fs.delete(new Path(dir, "missing.txt"), true));
assertTrue(fs.delete(dir, true));
}
Application-level acceptance test
Validate the complete Titan pipeline on ABFSS
The full Titan application-level E2E gate shown here uses Spark 3.5.2 + Hadoop 3.3.4 + SecureAzureBlobFileSystem over abfss://. The same DFS implementation was also validated against the Spark 4.0.0 / Hadoop 3.4.1 and Spark 4.1.0 / Hadoop 3.4.2 compatibility baselines listed above.
Titan repeatedly crosses the filesystem boundary while discovering source files, reading multiple formats, writing intermediate datasets, listing results and committing Spark output. The E2E run verifies that those contracts work together across the complete pipeline.
Source formats tested end to end
Nine input workflows. One ABFSS-backed Titan pipeline.
Reproduce the validation
Keep the compatibility claim executable
The useful outcome is not a one-off green run. The repository carries the Hadoop compatibility path so a future change can be checked against the same client contract.
./mvnw test
# final result
Tests run: 930
Failures: 0
Errors: 0
Skipped: 0
make test-java-compat
# final result
Tests run: 263
Failures: 0
Errors: 0
Skipped: 2
BUILD SUCCESS
Engineering takeaway
Move filesystem failures into the local feedback loop
The main benefit is not that Azure suddenly runs on a laptop. The benefit is that the same Spark and Hadoop filesystem path can now be exercised before a deployment reaches Azure, which moves protocol and integration failures much closer to the code change that caused them.
Real client behavior locally. Spark and Hadoop ABFS stay in the execution path instead of being replaced by storage mocks.
Earlier protocol feedback. Wrong DFS errors, listing behavior, metadata updates and rename semantics can fail in the local or CI loop.
Cleaner Azure validation. Cloud tests can focus on identity, private networking, service limits and production-specific behavior.
The result is a shorter feedback cycle without pretending that a local emulator is a substitute for Azure itself.
Shorter feedback loop
Catch the storage contract before cloud deployment
FAQ
Local ADLS Gen2, Hadoop ABFS and Spark questions
Practical answers about the protocol choices, test harness and scope of the Floci AZ contribution.
Can Apache Spark use ADLS Gen2 locally without a real Azure Storage Account?
Yes, provided the emulator reproduces the DFS wire semantics that Hadoop ABFS expects. The important part is not only accepting the same URLs; service-specific errors, ETags, resource types, listing behaviour and path update semantics must also match the client contract.
Why did BlobNotFound versus PathNotFound break FileSystem.delete()?
Hadoop ABFS interprets the ADLS service error code and JSON body. A Blob-style 404 is not semantically equivalent to a DFS PathNotFound response. With the correct DFS response, deleting a missing path maps to the normal Hadoop filesystem result: false.
Which Hadoop implementations were validated?
The repository compatibility suite directly uses hadoop-common 3.3.4 and hadoop-azure 3.3.4 with the real org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem class. Additional compatibility baselines cover Hadoop 3.4.1 with hadoop-azure 3.4.1 and Hadoop 3.4.2 with hadoop-azure 3.4.2 alongside Spark 4.0.0 and 4.1.0.
Was abfss:// tested too?
Yes. The Docker compatibility harness uses abfs:// over loopback HTTP so it can preserve the DFS host cleanly. External Spark validation used org.apache.hadoop.fs.azurebfs.SecureAzureBlobFileSystem and abfss:// paths.
Why is fs.azure.always.use.https=false set in the compatibility test?
In the Hadoop 3.3.4 repository compatibility harness, this setting defaults to true even for abfs://. That local harness deliberately forwards plaintext loopback traffic on port 80, so HTTPS is disabled only for that test. It does not alter Floci's runtime routing behaviour.
Does this implement complete ADLS Gen2 POSIX security?
No. Owner, group, permission and ACL metadata round-trip through the DFS/Hadoop wire protocol, but the emulator does not attempt to reproduce the complete Azure POSIX authorization engine.
How was the final implementation validated?
The final branch passed 930 repository tests and 263 Java compatibility tests with zero failures. The repository suite directly exercises the real Hadoop 3.3.4 ABFS client, the Titan application-level E2E gate runs Spark 3.5.2 over abfss:// across nine source workflows, and additional compatibility baselines validate the DFS implementation with Spark 4.0.0 / Hadoop 3.4.1 and Spark 4.1.0 / Hadoop 3.4.2.
Where is the open-source implementation?
The implementation was developed for the Floci AZ project and is available on the public feature branch linked from this article while the upstream contribution is reviewed.
Build closer to production
Make local data-platform tests exercise the real integration boundary
We design Spark, Databricks and Azure data platforms with repeatable infrastructure and test boundaries that cover more than application code alone.
A stronger test boundary