fabricQueryR helps you work with Microsoft Fabric directly from R.
You can use it to find Fabric workspaces and data items, query various Fabric data interfaces (SQL, DAX, KQL, GraphQL), work with OneLake files and tables, run Spark code, and start or monitor Fabric jobs.
Install the latest release from CRAN:
install.packages("fabricQueryR")Or install the development version from GitHub:
if (!requireNamespace("remotes", quietly = TRUE)) {
install.packages("remotes")
}
remotes::install_github("kennispunttwente/fabricQueryR")Below is a minimal example of how to find a workspace and some of its items.
To connect to Microsoft Fabric, you need a Microsoft Entra tenant ID and optionally a client ID. See the authentication vignette for more details on how to set these environment variables and sign in to Microsoft Fabric from R.
Once you know these values, you can discover workspaces and items like so:
library(fabricQueryR)
library(purrr)
# Always set your tenant ID:
Sys.setenv(FABRICQUERYR_TENANT_ID = "your-tenant-id")
# Optional, if your tenant does not permit the public Azure CLI client ID:
# Sys.setenv(FABRICQUERYR_CLIENT_ID = "your-app-client-id")
# Find all workspaces and select one by name
workspaces <- fabric_workspaces()
workspace <- workspaces |>
keep(~ .x$displayName == "ExampleWorkspace") |>
first()
# Find all items in the workspace
items <- fabric_items(workspace)
# Find the first item of each type
lakehouse <- fabric_lakehouses(workspace) |> first()
warehouse <- fabric_warehouses(workspace) |> first()
warehouse_snapshot <- fabric_warehouse_snapshots(workspace) |> first()
sql_database <- fabric_sql_databases(workspace) |> first()
semantic_model <- fabric_semantic_models(workspace) |> first()
kql_database <- fabric_kql_databases(workspace) |> first()
graphql_api <- fabric_graphql_apis(workspace) |> first()
notebook <- fabric_notebooks(workspace) |> first()In the next sections, you can see how to use these items to query data, run Spark code, and more.
Also see the reference for full function documentation and more examples.
Connect to a Warehouse, SQL Database, or a Lakehouse's SQL analytics endpoint. The result is a standard DBI connection, so it works with familiar R database packages such as DBI and dbplyr.
# Set up a DBI connection to a Lakehouse SQL endpoint
con <- fabric_sql_connect(lakehouse)
# List the tables
DBI::dbListTables(con)
# Run a SQL query and return the result as a tibble
df_sql <- DBI::dbGetQuery(
con,
"SELECT * FROM dbo.Customers WHERE region = 'West'"
)
# Close the connection when done
DBI::dbDisconnect(con)
# Or, run a single SQL query directly (without a connection object):
df_sql <- fabric_sql_query(
lakehouse,
"SELECT * FROM dbo.Customers WHERE region = 'West'"
)The default connection method requires
Microsoft ODBC Driver 18 for SQL Server.
An optional ADBC backend is available for Arrow-based workflows; see
fabric_sql_connect().
Run a DAX query against tables and measures in a Fabric or Power BI semantic model. The result is returned as a tibble.
df_dax <- fabric_pbi_dax_query(
connstr = semantic_model,
dax = "EVALUATE TOPN(1000, 'Customers')"
)
# Use the newer typed Arrow API on modern semantic models
arrow_stream <- fabric_pbi_dax_query(
connstr = semantic_model,
dax = "EVALUATE TOPN(1000, 'Customers')",
api = "arrow",
result = "arrow_stream"
)
reader <- arrow::as_record_batch_reader(arrow_stream)The signed-in account needs access to the workspace or Read and Build permissions on the semantic model.
Run SparkR, PySpark, Scala, or Spark SQL remotely in Fabric, and get the result in your local R session.
# Run a single SparkR statement through Livy
livy_sparkr_result <- fabric_livy_query(
livy_url = lakehouse,
kind = "sparkr",
code = "print(1 + 2)"
)
# Run a reusable session, with deterministic cleanup on errors
run_shared_state <- function(lakehouse) {
livy <- fabric_livy_session(lakehouse)
on.exit(livy$close(), add = TRUE)
livy$wait()
livy$run("shared_value = 40", kind = "pyspark")
livy$run("print(shared_value + 2)", kind = "pyspark")
}
# Let Fabric pack isolated REPLs into high-concurrency Spark sessions
run_high_concurrency <- function(lakehouse) {
livy <- fabric_livy_session(
lakehouse,
high_concurrency = TRUE,
session_tag = "report-workers"
)
on.exit(livy$close(), add = TRUE)
livy$wait()
livy$run("SELECT current_timestamp()", kind = "sql")
}
# Submit an application file as an independent batch job
batch <- fabric_livy_batch_submit(
lakehouse,
file = paste0(
"abfss://workspace@onelake.dfs.fabric.microsoft.com/",
"lakehouse/Files/jobs/daily.py"
),
wait = TRUE,
cancel_on_timeout = TRUE
)
batch$result()Copy the session or batch connection URL from Lakehouse settings > Livy
endpoint, or pass a Lakehouse returned by fabric_lakehouses() as above.
Read a Delta table stored in OneLake and return its current rows as a tibble.
df_onelake <- fabric_onelake_read_delta_table(
table_path = "Customers",
workspace_name = workspace,
lakehouse_name = lakehouse
)
# Optional: return an Arrow C stream instead of a tibble
stream <- fabric_onelake_read_delta_table(
table_path = "Customers",
workspace_name = workspace,
lakehouse_name = lakehouse,
result = "arrow_stream"
)
reader <- arrow::as_record_batch_reader(stream)This direct reader intentionally stops on Delta features unsupported by its
pinned deltalake runtime, including type widening, V2 checkpoints, and
Variant. Deletion-vector tables are supported by the pinned reader. Query
unsupported tables through fabric_sql_query() or Spark through the Livy
helpers instead. See Microsoft's
runtime compatibility guidance.
List, inspect, download, upload, or delete files in OneLake.
Paths in a Lakehouse usually start with Files/.
files <- fabric_onelake_list(
workspace,
lakehouse,
path = "Files/fixtures",
recursive = TRUE
)
fabric_onelake_download(
workspace,
lakehouse,
path = "Files/fixtures/basic.csv",
dest = "basic.csv"
)Run a KQL query against a KQL database in an Eventhouse. A single result table is returned as a tibble, with Kusto data types converted to R data types.
df_kql <- fabric_kql_query(
kql_database,
query = paste(
"declare query_parameters(selected_type:string);",
"Events | where EventType == selected_type | take 100"
),
parameters = list(selected_type = "Warning")
)Call an API for GraphQL item that has already been configured in Fabric.
The result keeps the nested GraphQL data and any GraphQL-level errors separate.
graphql_result <- fabric_graphql_query(
graphql_api,
query = paste(
"query Customers($region: String!) {",
" customers(filter: {region: {eq: $region}}) {",
" items { id name region }",
" }",
"}"
),
variables = list(region = "West"),
operation_name = "Customers"
)
graphql_result$data$customers$items
graphql_result$errorsStart a notebook, data pipeline, or Spark job definition, then wait for completion or inspect/cancel it from R.
job <- fabric_job_run(
notebook,
parameters = list(mode = "incremental", batch_size = 100L),
default_lakehouse = lakehouse
)
result <- fabric_job_wait(job, timeout = 900)
result$status
result$exit_valueMicrosoft Fabric brings data engineering, data warehousing, real-time analytics, and Power BI together in one platform, with OneLake as its shared storage layer.
When our organization started using Fabric, accessing its data from R was not yet straightforward. This package grew from the helper functions created to make that work easier and more consistent for other R users.
I (Luka Koning) am no longer associated with Kennispunt Twente. I maintain this open-source R package in my personal capacity. The repository remains under the Kennispunt Twente GitHub organization for historical reasons, and because the early development was done while I was affiliated with them.