Reference Documentation & Whitepapers
Self-contained, production-grade technical whitepapers covering Apache Spark execution internals, Remote Shuffle Manager lifecycles, Runtime 2.0 migration recipes, and Polaris distributed SQL architecture.
Fabric Spark Engine Internals
Execution Model, Catalyst, Tungsten, AQE, NEE, Memory & Capacity Diagnostics
Comprehensive deep dive into the Apache Spark engine in Microsoft Fabric. Traces the full query execution pipeline, memory allocator boundaries, native execution engine fallbacks, and SKU throttling mechanics.
- •AQE runtime plan adaptation & shuffle partition coalesce
- •Native Execution Engine (NEE) C++ vectorization boundary & fallback matrix
- •Unified Memory Manager: Execution vs Storage vs Off-Heap allocations
Efficient Scaledown & Remote Shuffle Manager
RSM, Shuffle Migration, Decommission Lifecycle & Enterprise Guidance
Practitioner architectural whitepaper tracing Remote Shuffle Manager (RSM), shuffle data lifecycle across node decommissions, and dynamic scale-down optimization in Fabric Spark.
- •SPARK-20624 decommission lifecycle & BlockManager block migration
- •SortShuffleManager vs Remote Shuffle Manager storage topology
- •ShuffleDataIO abstractions and AQE shuffle write interactions
Fabric Runtime 2.0 Deep Dive
Apache Spark 4.1, Delta 4.2 & Migration Architecture
Detailed migration and architecture guide for Fabric Runtime 2.0. Analyzes the Spark 4.x ANSI mode default, Native Execution Engine behavior, %%configure precedence, and production recipes.
- •Spark 4.1 ANSI mode default × NEE fallback interaction matrix
- •V-Order optimization decision tree and write latency tradeoffs
- •%%configure scope necessity across notebook vs pipeline execution
OneLake Storage, Polaris & Direct Lake
Shortcut Resolution, Distributed SQL Compilation & Direct Lake Guardrails
Architectural deep dive covering OneLake hierarchical storage, shortcut identity resolution, Polaris distributed SQL compilation, and Direct Lake memory paging.
- •OneLake shortcut identity resolution & delegation flow
- •Polaris distributed SQL compilation pipeline: frontend to execution
- •Direct Lake file fragmentation diagnostics and memory guardrails
Fabric Deep Dives — SQL, Functions & dbt
Spark View Types, Fabric User Data Functions & dbt-on-Fabric
Synthesized practitioner guide covering Spark view and function types with executed test suites, User Data Functions (UDFs), Fabric SQL Database access paths, and dbt adapter topologies.
- •Executed VERIFIED / FALSIFIED matrix of Spark view & function types
- •Native Execution Engine do's and don'ts practice sheet
- •Three distinct Fabric SQL Database notebook access paths