First-Party Deep Dives & Engine Internals

Reference Documentation & Whitepapers

Self-contained, production-grade technical whitepapers covering Apache Spark execution internals, Remote Shuffle Manager lifecycles, Runtime 2.0 migration recipes, and Polaris distributed SQL architecture.

Interactive
spark
capacity
~45 min

Fabric Spark Engine Internals

Execution Model, Catalyst, Tungsten, AQE, NEE, Memory & Capacity Diagnostics

Comprehensive deep dive into the Apache Spark engine in Microsoft Fabric. Traces the full query execution pipeline, memory allocator boundaries, native execution engine fallbacks, and SKU throttling mechanics.

Core Insights
  • AQE runtime plan adaptation & shuffle partition coalesce
  • Native Execution Engine (NEE) C++ vectorization boundary & fallback matrix
  • Unified Memory Manager: Execution vs Storage vs Off-Heap allocations
spark
fabric platform
~20 min

Efficient Scaledown & Remote Shuffle Manager

RSM, Shuffle Migration, Decommission Lifecycle & Enterprise Guidance

Practitioner architectural whitepaper tracing Remote Shuffle Manager (RSM), shuffle data lifecycle across node decommissions, and dynamic scale-down optimization in Fabric Spark.

Core Insights
  • SPARK-20624 decommission lifecycle & BlockManager block migration
  • SortShuffleManager vs Remote Shuffle Manager storage topology
  • ShuffleDataIO abstractions and AQE shuffle write interactions
spark
lakehouse
~18 min

Fabric Runtime 2.0 Deep Dive

Apache Spark 4.1, Delta 4.2 & Migration Architecture

Detailed migration and architecture guide for Fabric Runtime 2.0. Analyzes the Spark 4.x ANSI mode default, Native Execution Engine behavior, %%configure precedence, and production recipes.

Core Insights
  • Spark 4.1 ANSI mode default × NEE fallback interaction matrix
  • V-Order optimization decision tree and write latency tradeoffs
  • %%configure scope necessity across notebook vs pipeline execution
onelake
polaris
direct lake
sql database
~16 min

OneLake Storage, Polaris & Direct Lake

Shortcut Resolution, Distributed SQL Compilation & Direct Lake Guardrails

Architectural deep dive covering OneLake hierarchical storage, shortcut identity resolution, Polaris distributed SQL compilation, and Direct Lake memory paging.

Core Insights
  • OneLake shortcut identity resolution & delegation flow
  • Polaris distributed SQL compilation pipeline: frontend to execution
  • Direct Lake file fragmentation diagnostics and memory guardrails
sql database
spark
fabric data-agent
~22 min

Fabric Deep Dives — SQL, Functions & dbt

Spark View Types, Fabric User Data Functions & dbt-on-Fabric

Synthesized practitioner guide covering Spark view and function types with executed test suites, User Data Functions (UDFs), Fabric SQL Database access paths, and dbt adapter topologies.

Core Insights
  • Executed VERIFIED / FALSIFIED matrix of Spark view & function types
  • Native Execution Engine do's and don'ts practice sheet
  • Three distinct Fabric SQL Database notebook access paths