Parquet Folder logo

Columnar files New

Connect Parquet Folder to Excel, Sheets and AI

Query a folder of Parquet files as one table. Parquet carries its own column types, so nothing here is guessed — a decimal stays a decimal and a timestamp stays a timestamp.

1connection
0inbound ports
read-onlyenforced

One connection, every surface

Where your Parquet Folder data can go

Connect Parquet Folder once and the same read-only connection feeds all of these — no second setup, no second copy of the data.

Supported

Parquet Folder to Excel

Microsoft Excel · Excel add-in

Pull live Parquet Folder results straight into a worksheet and refresh them on demand — desktop Excel, Excel Online, Microsoft 365.

How Excel works no Parquet Folder walkthrough written yet
Supported

Parquet Folder to Google Sheets

Sheets add-on

Run a saved Parquet Folder query from the sidebar and drop the rows into the sheet. Shared collaborators can refresh it themselves.

How Google Sheets works no Parquet Folder walkthrough written yet
Supported

Parquet Folder MCP server

Claude, Cursor and MCP clients

Give an AI assistant read-only access to Parquet Folder with the schema it needs to write correct SQL — no credentials in the chat.

How MCP works no Parquet Folder walkthrough written yet
Supported

Parquet Folder REST API

HTTP endpoint

Publish a Parquet Folder query as an authenticated JSON endpoint any application can call, with an OpenAPI 3.1 spec and ready-made Postman, Insomnia and Hoppscotch collections. No database port is opened.

How REST API works no Parquet Folder walkthrough written yet
Supported

Parquet Folder to Airtable

Automation platform

Sync Parquet Folder rows into an Airtable base on a schedule, or fetch them inside an Airtable automation script.

How Airtable works no Parquet Folder walkthrough written yet
Supported

Parquet Folder to Baserow

Automation platform

Feed a Baserow table from Parquet Folder over the REST endpoint — self-hosted or Baserow cloud.

How Baserow works no Parquet Folder walkthrough written yet
Supported

Parquet Folder to SeaTable

Automation platform

Keep a SeaTable base current with Parquet Folder data without exporting a file or exposing the database.

How SeaTable works no Parquet Folder walkthrough written yet
Supported

Parquet Folder to Smartsheet

Automation platform

Push Parquet Folder results into a Smartsheet grid so plans and reports read from the source system, not last week's export.

How Smartsheet works no Parquet Folder walkthrough written yet
Supported

Parquet Folder to Anvil

Anvil Works · App platform

Back an Anvil Python app with Parquet Folder through the REST endpoint instead of embedding database credentials in the app.

How Anvil works no Parquet Folder walkthrough written yet
Supported

Parquet Folder to Power BI

Power Query M

Paste the generated Power Query M into the Power BI Advanced Editor and the report reads live Parquet Folder results over HTTPS — no ODBC driver, no database port opened.

How Power BI works no Parquet Folder walkthrough written yet
Supported

Parquet Folder alerts and reports

Slack · Discord · Email · Webhook

Put a Parquet Folder query on a schedule and have the rows delivered to Slack, Discord, email or a signed webhook — or hold the message until a row count, threshold or percentage change crosses the line you set.

How alerts and reports work no Parquet Folder walkthrough written yet

How it works

5 steps, no inbound firewall change

01

Install the Network Agent where it can see the folder — an export directory, a small data lake, a vendor drop.

02

Point the connector at the folder or the whole tree; every .parquet file is matched.

03

Parquet stores its schema inside each file, so the agent reads those descriptions rather than sampling rows. No type is inferred from text, and no sampling error is possible.

04

The files load into one table, stacked, with a _source_file column on every row naming its file.

05

Query it with the same SQL as your databases — and join it to them in the same statement.

Feature deep-dive

What Parquet Folder gives you

Why this one is exact

  • Types come from the file itself — decimals, timestamps and integers arrive as what they already are.
  • Column union — if some files carry a column the others lack, it simply reads as empty for those files instead of breaking the set.
  • If two files genuinely disagree on a column's type, the connector refuses to pin rather than flattening both to text. You are told which file and which column.
-- Rows and revenue per export file
SELECT   _source_file, COUNT(*) AS rows, SUM(amount) AS total
FROM     data
GROUP BY _source_file

After the first sync

A later file that adds a column, drops a pinned one, or changes a type is held back and recorded in files_events by name. The rest of the folder keeps loading, so one malformed export never takes the table offline. Reading is cheap: only files that actually changed are re-read.

Shared by the whole File Set family

Every folder connector also gives you

  • files_current — a live inventory: every file the connector can see right now, with its path, size and modified date.
  • files_events — the audit trail: what appeared, what changed, what vanished, and anything held back, with the filename and the reason.
  • directories and volumes — per-folder totals, daily growth history, and how much room is left on the drive.
  • Only files that actually changed are re-read on each sync, so a folder of 50,000 files is not re-parsed because one new export landed.

Good to know

  • Read-only, enforced — one statement at a time, SELECT and friends only. Your files are never written, moved or renamed.
  • Nothing is uploaded — the data is cached, encrypted, on your own machine beside the agent. Only the result of a query leaves your network.
  • Sensible defaults — up to 250,000 files, 16 folders deep, 512 MB per file, all adjustable. Recycle bins, .git and node_modules are always skipped.
  • The SQL dialect is DuckDB — the same SQL you would write against any other connector.
  • Requires Network Agent 2.6 or newer.

Cross-source SQL

Join Parquet Folder to the rest of your data

A folder of files is a set of SQL tables like any other, so one statement can join it to a database and an API at once. Each source runs only the part it can, streams the result back, and the join happens centrally — the sources never talk to each other and nothing is copied anywhere.

3 connections · 3 agents

Parquet Folder Columnar files
PostgreSQL Relational engine
Stripe Payments & billing

One statement

-- nothing copied, nothing merged, nothing scheduled
SELECT   c.region, COUNT(*) AS orders, SUM(i.amount_due) AS invoiced
FROM     parquet_lake.fileset.data1 f
JOIN     pg_crm.public.customers2   c ON c.id = f.customer_id
JOIN     billing.stripe.invoices3   i ON i.customer = c.stripe_id
GROUP BY c.region
ORDER BY invoiced DESC;

The three parts are connection, schema and table — and the connection name is whatever you called it. Illustrative columns; your tables will be your tables. Read-only applies to every piece: SELECT, WITH and EXPLAIN only, with a ceiling on how much any one source may hand over for a single query. How federated queries work

Connection details

What Parquet Folder needs

Folder
One or more roots — local disk, mapped drive or UNC share; every .parquet file is matched, subfolders included
Schema
Read from each file's own Parquet footer — no sampling, no inference, no type worked out from text
Column union
A column only some files carry reads as empty in the others rather than breaking the set
Type conflicts
If two files genuinely disagree on a column's type the connector refuses to pin and names the file and the column, instead of flattening both to text
Tables
One table, every file stacked into it, each row carrying a _source_file column naming its file
Limits
Up to 250,000 files, 16 folder levels deep, 512 MB per file — all adjustable
Credentials
None — there is no server; the agent reads the files in place
SQL dialect
DuckDB — standard SQL, nothing folder-specific to learn
Agent
Network Agent 2.6 or newer

There is no database server in this picture. The agent reads the .parquet files where they already live, caches the rows in an encrypted DuckDB store on the same machine, and re-reads only files that actually changed — a folder of 50,000 files is not re-parsed because one new export landed. Nothing is uploaded to Query Streams; the only thing that ever leaves your network is the result of a query.

Parquet is the one format in this family where nothing has to be inferred. Every other folder connector works out what a column is by looking at values; Parquet already knows, so the agent reads the description rather than the data. That makes setup cheap even on very large folders — only the footers are examined — and it removes the whole class of problem where a column looks numeric for the first thousand rows and then turns out not to be.

Vendor documentation: parquet.apache.org

FAQ

Questions about Parquet Folder

Which tools can read Parquet Folder data through Query Streams?

All of them, from one connection: Excel, Google Sheets, MCP, REST API, Airtable, Baserow, SeaTable, Smartsheet, Anvil, Power BI, scheduled alerts and reports. Connect the folder once and every surface reads the same read-only connection — there is no per-tool setup and no second copy of the data.

Do my files get uploaded to Query Streams?

No. The Network Agent reads the files in place and caches rows in an encrypted store on the same machine. The files themselves never leave your network — only the result rows of a query do, over a single outbound encrypted connection with no inbound firewall port.

Can Query Streams change, move or rename my files?

No. Files are opened strictly read-only and are never written, moved or renamed. Queries are enforced read-only at the point of execution — one statement at a time, SELECT and friends only.

What does Query Streams need to connect to Parquet Folder?

A folder path the agent machine can see — no server, no credentials, no drivers to install. Folder: One or more roots — local disk, mapped drive or UNC share; every .parquet file is matched, subfolders included. Schema: Read from each file's own Parquet footer — no sampling, no inference, no type worked out from text. Column union: A column only some files carry reads as empty in the others rather than breaking the set. Type conflicts: If two files genuinely disagree on a column's type the connector refuses to pin and names the file and the column, instead of flattening both to text.

Can I join a folder of files to a database in the same query?

Yes — that is a federated query. One statement can reference Parquet Folder and your other connections at once, written as connection.schema.table. Each source runs only the part it can and streams the result back; the join happens centrally, so the sources never connect to each other and nothing is copied or scheduled. Read-only applies to every piece — SELECT, WITH and EXPLAIN only — and there is a ceiling on how much any one source may hand over for a single query. Federated queries are a plan feature; the federated queries page carries the current source and size limits.

Put Parquet Folder where the work happens

Install the agent, point it at your folder, and pick a destination.

Read-only Outbound only Credentials stay on the agent