AptlyStar

Databricks

Run SQL queries and manage jobs on Databricks

使用说明

Connect to Databricks to execute SQL queries against SQL warehouses, trigger and monitor job runs, manage clusters, and retrieve run outputs. Requires a Personal Access Token and workspace host URL.

工具

databricks_execute_sql

Execute a SQL statement against a Databricks SQL warehouse and return results inline. Supports parameterized queries and Unity Catalog.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token
warehouseIdstring是The ID of the SQL warehouse to execute against
statementstring是The SQL statement to execute (max 16 MiB)
catalogstring否Unity Catalog name (equivalent to USE CATALOG)
schemastring否Schema name (equivalent to USE SCHEMA)
rowLimitnumber否Maximum number of rows to return
waitTimeoutstring否How long to wait for results (e.g., "50s"). Range: "0s" or "5s" to "50s". Default: "50s"

输出

参数类型描述
statementIdstringUnique identifier for the executed statement
statusstringExecution status (SUCCEEDED, PENDING, RUNNING, FAILED, CANCELED, CLOSED)
columnsarrayColumn schema of the result set
↳ namestringColumn name
↳ positionnumberColumn position (0-based)
↳ typeNamestringColumn type (STRING, INT, LONG, DOUBLE, BOOLEAN, TIMESTAMP, DATE, DECIMAL, etc.)
dataarrayResult rows as a 2D array of strings where each inner array is a row of column values
totalRowsnumberTotal number of rows in the result
truncatedbooleanWhether the result set was truncated due to row_limit or byte_limit

databricks_get_statement

Poll a SQL statement by its ID to retrieve status and results. Use this after Execute SQL when a query runs longer than the wait timeout.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token
statementIdstring是The ID of the statement to fetch (returned by Execute SQL)

输出

参数类型描述
statementIdstringUnique identifier for the statement
statusstringExecution status (SUCCEEDED, PENDING, RUNNING, FAILED, CANCELED, CLOSED)
columnsarrayColumn schema of the result set
↳ namestringColumn name
↳ positionnumberColumn position (0-based)
↳ typeNamestringColumn type (STRING, INT, LONG, DOUBLE, BOOLEAN, TIMESTAMP, DATE, DECIMAL, etc.)
dataarrayResult rows as a 2D array of strings where each inner array is a row of column values
totalRowsnumberTotal number of rows in the result
truncatedbooleanWhether the result set was truncated due to row_limit or byte_limit

databricks_list_warehouses

List all SQL warehouses in a Databricks workspace including their size, state, and type. Use this to discover the warehouse ID needed for Execute SQL.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token

输出

参数类型描述
warehousesarrayList of SQL warehouses in the workspace
↳ warehouseIdstringUnique warehouse identifier
↳ namestringWarehouse display name
↳ clusterSizestringWarehouse size (e.g., 2X-Small, Small, Medium, Large)
↳ statestringCurrent state (STARTING, RUNNING, STOPPING, STOPPED, DELETING, DELETED)
↳ warehouseTypestringWarehouse type (CLASSIC, PRO)
↳ creatorNamestringEmail of the warehouse creator
↳ autoStopMinutesnumberMinutes of inactivity before auto-stop (0 = disabled)
↳ numClustersnumberCurrent number of running clusters
↳ minNumClustersnumberMinimum cluster count for scaling
↳ maxNumClustersnumberMaximum cluster count for scaling
↳ numActiveSessionsnumberNumber of active sessions
↳ enableServerlessComputebooleanWhether serverless compute is enabled

databricks_list_jobs

List all jobs in a Databricks workspace with optional filtering by name.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token
limitnumber否Maximum number of jobs to return (range 1-100, default 20)
offsetnumber否Offset for pagination
namestring否Filter jobs by exact name (case-insensitive)
expandTasksboolean否Include task and cluster details in the response (max 100 elements)

输出

参数类型描述
jobsarrayList of jobs in the workspace
↳ jobIdnumberUnique job identifier
↳ namestringJob name
↳ createdTimenumberJob creation timestamp (epoch ms)
↳ creatorUserNamestringEmail of the job creator
↳ maxConcurrentRunsnumberMaximum number of concurrent runs
↳ formatstringJob format (SINGLE_TASK or MULTI_TASK)
hasMorebooleanWhether more jobs are available for pagination
nextPageTokenstringToken for fetching the next page of results

databricks_get_job

Get the full definition and settings of a single Databricks job by its job ID.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token
jobIdnumber是The canonical identifier of the job to retrieve

输出

参数类型描述
jobIdnumberThe job ID
namestringJob name
creatorUserNamestringEmail of the job creator
runAsUserNamestringUser the job runs as
createdTimenumberJob creation timestamp (epoch ms)
formatstringJob format (SINGLE_TASK or MULTI_TASK)
maxConcurrentRunsnumberMaximum number of concurrent runs
timeoutSecondsnumberJob-level timeout in seconds (0 or null means no timeout)
scheduleobjectCron schedule configuration (quartz_cron_expression, timezone_id, pause_status)
tagsobjectKey-value tags applied to the job
tasksarrayTask definitions for the job (empty for single-task jobs)

databricks_run_job

Trigger an existing Databricks job to run immediately with optional job-level or notebook parameters.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token
jobIdnumber是The ID of the job to trigger
jobParametersstring否Job-level parameter overrides as a JSON object (e.g., {"key": "value"})
notebookParamsstring否Notebook task parameters as a JSON object (e.g., {"param1": "value1"})
idempotencyTokenstring否Idempotency token to prevent duplicate runs (max 64 characters)

输出

参数类型描述
runIdnumberThe globally unique ID of the triggered run
numberInJobnumberThe sequence number of this run among all runs of the job

databricks_get_run

Get the status, timing, and details of a Databricks job run by its run ID.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token
runIdnumber是The canonical identifier of the run
includeHistoryboolean否Include repair history in the response
includeResolvedValuesboolean否Include resolved parameter values in the response

输出

参数类型描述
runIdnumberThe run ID
jobIdnumberThe job ID this run belongs to
runNamestringName of the run
runTypestringType of run (JOB_RUN, WORKFLOW_RUN, SUBMIT_RUN)
attemptNumbernumberRetry attempt number (0 for initial attempt)
stateobjectRun state information
↳ lifeCycleStatestringLifecycle state (QUEUED, PENDING, RUNNING, TERMINATING, TERMINATED, SKIPPED, INTERNAL_ERROR, BLOCKED, WAITING_FOR_RETRY)
↳ resultStatestringResult state (SUCCESS, FAILED, TIMEDOUT, CANCELED, SUCCESS_WITH_FAILURES, UPSTREAM_FAILED, UPSTREAM_CANCELED, EXCLUDED)
↳ stateMessagestringDescriptive message for the current state
↳ userCancelledOrTimedoutbooleanWhether the run was cancelled by user or timed out
startTimenumberRun start timestamp (epoch ms)
endTimenumberRun end timestamp (epoch ms, 0 if still running)
setupDurationnumberCluster setup duration (ms)
executionDurationnumberExecution duration (ms)
cleanupDurationnumberCleanup duration (ms)
queueDurationnumberTime spent in queue before execution (ms)
runPageUrlstringURL to the run detail page in Databricks UI
creatorUserNamestringEmail of the user who triggered the run

databricks_list_runs

List job runs in a Databricks workspace with optional filtering by job, status, and time range.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token
jobIdnumber否Filter runs by job ID. Omit to list runs across all jobs
activeOnlyboolean否Only include active runs (PENDING, RUNNING, or TERMINATING)
completedOnlyboolean否Only include completed runs
limitnumber否Maximum number of runs to return (range 1-24, default 20)
offsetnumber否Offset for pagination
runTypestring否Filter by run type (JOB_RUN, WORKFLOW_RUN, SUBMIT_RUN)
startTimeFromnumber否Filter runs started at or after this timestamp (epoch ms)
startTimeTonumber否Filter runs started at or before this timestamp (epoch ms)

输出

参数类型描述
runsarrayList of job runs
↳ runIdnumberUnique run identifier
↳ jobIdnumberJob this run belongs to
↳ runNamestringRun name
↳ runTypestringRun type (JOB_RUN, WORKFLOW_RUN, SUBMIT_RUN)
↳ stateobjectRun state information
↳ lifeCycleStatestringLifecycle state (QUEUED, PENDING, RUNNING, TERMINATING, TERMINATED, SKIPPED, INTERNAL_ERROR, BLOCKED, WAITING_FOR_RETRY)
↳ resultStatestringResult state (SUCCESS, FAILED, TIMEDOUT, CANCELED, SUCCESS_WITH_FAILURES, UPSTREAM_FAILED, UPSTREAM_CANCELED, EXCLUDED)
↳ stateMessagestringDescriptive state message
↳ userCancelledOrTimedoutbooleanWhether the run was cancelled by user or timed out
↳ startTimenumberRun start timestamp (epoch ms)
↳ endTimenumberRun end timestamp (epoch ms)
hasMorebooleanWhether more runs are available for pagination
nextPageTokenstringToken for fetching the next page of results

databricks_cancel_run

Cancel a running or pending Databricks job run. Cancellation is asynchronous; poll the run status to confirm termination.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token
runIdnumber是The canonical identifier of the run to cancel

输出

参数类型描述
successbooleanWhether the cancel request was accepted

databricks_get_run_output

Get the output of a completed Databricks job run, including notebook results, error messages, and logs. For multi-task jobs, use the task run ID (not the parent run ID).

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token
runIdnumber是The run ID to get output for. For multi-task jobs, use the task run ID

输出

参数类型描述
notebookOutputobjectNotebook task output (from dbutils.notebook.exit())
↳ resultstringValue passed to dbutils.notebook.exit() (max 5 MB)
↳ truncatedbooleanWhether the result was truncated
errorstringError message if the run failed or output is unavailable
errorTracestringError stack trace if available
logsstringLog output (last 5 MB) from spark_jar, spark_python, or python_wheel tasks
logsTruncatedbooleanWhether the log output was truncated

databricks_list_clusters

List all clusters in a Databricks workspace including their state, configuration, and resource details.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token

输出

参数类型描述
clustersarrayList of clusters in the workspace
↳ clusterIdstringUnique cluster identifier
↳ clusterNamestringCluster display name
↳ statestringCurrent state (PENDING, RUNNING, RESTARTING, RESIZING, TERMINATING, TERMINATED, ERROR, UNKNOWN)
↳ stateMessagestringHuman-readable state description
↳ creatorUserNamestringEmail of the cluster creator
↳ sparkVersionstringSpark runtime version (e.g., 13.3.x-scala2.12)
↳ nodeTypeIdstringWorker node type identifier
↳ driverNodeTypeIdstringDriver node type identifier
↳ numWorkersnumberNumber of worker nodes (for fixed-size clusters)
↳ autoscaleobjectAutoscaling configuration (null for fixed-size clusters)
↳ minWorkersnumberMinimum number of workers
↳ maxWorkersnumberMaximum number of workers
↳ clusterSourcestringOrigin (API, UI, JOB, MODELS, PIPELINE, PIPELINE_MAINTENANCE, SQL)
↳ autoterminationMinutesnumberMinutes of inactivity before auto-termination (0 = disabled)
↳ startTimenumberCluster start timestamp (epoch ms)

databricks_get_cluster

Get the state, configuration, and resource details of a single Databricks cluster by its cluster ID.

输入

参数类型必填描述
hoststring是Databricks workspace host (e.g., dbc-abc123.cloud.databricks.com)
apiKeystring是Databricks Personal Access Token
clusterIdstring是The ID of the cluster to retrieve

输出

参数类型描述
clusterobjectCluster detail
↳ clusterIdstringUnique cluster identifier
↳ clusterNamestringCluster display name
↳ statestringCurrent state (PENDING, RUNNING, RESTARTING, RESIZING, TERMINATING, TERMINATED, ERROR, UNKNOWN)
↳ stateMessagestringHuman-readable state description
↳ creatorUserNamestringEmail of the cluster creator
↳ sparkVersionstringSpark runtime version (e.g., 13.3.x-scala2.12)
↳ nodeTypeIdstringWorker node type identifier
↳ driverNodeTypeIdstringDriver node type identifier
↳ numWorkersnumberNumber of worker nodes (for fixed-size clusters)
↳ autoscaleobjectAutoscaling configuration (null for fixed-size clusters)
↳ minWorkersnumberMinimum number of workers
↳ maxWorkersnumberMaximum number of workers
↳ clusterSourcestringOrigin (API, UI, JOB, MODELS, PIPELINE, PIPELINE_MAINTENANCE, SQL)
↳ autoterminationMinutesnumberMinutes of inactivity before auto-termination (0 = disabled)
↳ startTimenumberCluster start timestamp (epoch ms)

On this page