S3
Parquet
Privacy
DuckDB
Cloud Storage
Featured

How to Open Parquet from S3 Privately in Your Browser

Query Parquet on AWS S3, Cloudflare R2, GCS, or MinIO without proxying object data through viewparquet. Compatible endpoints can use range reads, and credentials stay in the browser.

By 11 min read

Most Parquet workflows eventually land in object storage — analytics buckets on AWS S3, exports on Cloudflare R2, Google Cloud Storage interoperability endpoints, or a MinIO cluster behind the firewall. The usual next step is a notebook or CLI command just to inspect a schema or count rows. viewparquet can point DuckDB-WASM at the object from the browser, run SQL, and inspect results without proxying the object data through a viewparquet data server.

This guide starts with a reproducible public file, then covers private credentials, provider settings, CORS, range requests, presigned URLs, and the failure modes that matter in a browser.

Reproduce the public remote-read path

Use viewparquet's own checked-in sample before debugging your bucket:

  1. Open the remote viewer.
  2. Choose Public or presigned.
  3. Paste this complete URL:
notes.txt
https://viewparquet.com/data/flights-200k.parquet
  1. Leave File type on Auto-detect and select Load data.
  2. In the SQL editor, run:
query.sql
SELECT
  count(*) AS rows,
  min(delay) AS min_delay,
  max(delay) AS max_delay
FROM data;

The repository copy used for this release contains 231,083 rows with delay (SMALLINT), distance (SMALLINT), and time (FLOAT) columns. It is 1,129,931 bytes, begins and ends with the Parquet PAR1 marker, and has SHA-256 6057fc59877fa24f0b9ece1e1807ec2709348cc7e687a871293a96b1040f4976 as checked on August 19, 2026.

That check proves the public-URL dialog, Parquet reader, and SQL path against one controlled file. It is not a speed benchmark, a promise that every host permits browser reads, or proof that a third-party server honors byte ranges.

What stays private — and what still crosses the network

For viewing and SQL, viewparquet does not proxy or receive the remote object data. The browser requests the configured storage host directly.

  • Access keys and session tokens are registered as a scoped DuckDB secret inside the current browser session. They are not saved with Recent sources or sent to viewparquet servers.
  • Presigned URLs contain temporary authorization in their query string. viewparquet keeps URLs with query parameters in session storage rather than durable local storage, but they remain bearer credentials. Clear Recent sources before sharing the browser.
  • Remote object bytes come from the object host. In DevTools → Network, confirm that the data request goes to S3, R2, GCS, MinIO, or the HTTPS host you entered.
  • Optional AI is separate. If invoked, it sends your messages and configured context to the provider you choose; chat tools and saved AI context may also provide structural metadata and SQL.
  • The page can make other disclosed asset, telemetry, error-monitoring, support, or AI requests. Aggregate product analytics records categories and outcomes rather than the remote URL, credentials, SQL text, cell values, or exported contents; inspect the Network panel and privacy page for the other disclosed services.

Path 1: Public HTTPS or a presigned URL

Use the Public or presigned tab for an object already reachable over HTTPS.

  1. Paste the full URL, including every query parameter on a presigned URL.
  2. Choose the file type explicitly if the path has no recognizable extension.
  3. Load the object, then begin with a narrow query.
  4. Do not enter separate access keys for a presigned URL; the URL already carries temporary authorization.

A presigned URL can stop working because it expired, its signing credentials expired, the wrong region or HTTP method was signed, the object moved, CORS blocks the browser, or a proxy rewrote the URL. Generate a fresh URL and paste it without decoding, shortening, or editing the query string.

A presigned URL authorizes the object request; it does not bypass browser CORS. The bucket still has to allow the page origin and required request headers.

Path 2: A public s3:// object

For an anonymously readable bucket, use the URL tab with an object URI:

notes.txt
s3://my-public-bucket/path/to/file.parquet

Anonymous S3 reads default to us-east-1 in the current flow. If the bucket is private, the region differs, or the anonymous path fails, use an HTTPS URL or the private-bucket path below.

Path 3: Private AWS, R2, GCS, or MinIO

Choose Private bucket and enter an s3://bucket/key.parquet path. The credential form creates a secret scoped to that bucket so it is not applied to unrelated reads.

ProviderCredentialsEndpoint and defaultsFrequent mistake
Amazon S3Access key + secret; optional STS tokenLeave endpoint blank; set the actual bucket region; virtual-hosted style by defaultWrong region, missing s3:GetObject, or an STS token entered without its session token
Cloudflare R2R2 S3 access key + secret<account_id>.r2.cloudflarestorage.com; region auto; path-stylePasting https:// into the endpoint or using an API token instead of an R2 S3 key
Google Cloud StorageHMAC interoperability access ID + secretstorage.googleapis.com; region auto; path-styleUsing OAuth/service-account JSON instead of an HMAC key
MinIO / customThe endpoint's S3-compatible key + secretHost and optional port; path-style by default; region must match the serverUsing HTTP from the hosted HTTPS app, a self-signed certificate, or the wrong path/vhost style

For MinIO or another custom store, the hosted https://viewparquet.com page normally needs an HTTPS endpoint with a browser-trusted certificate. Browsers can block an HTTP object request as mixed content; an HTTP endpoint is mainly useful in a permitted local-development setup.

Keys and tokens live in the form state and DuckDB instance for the current session. Reopening a private Recent source restores connection hints, not secrets, and prompts for credentials again.

A minimal browser CORS starting point

For an AWS-style CORS editor, this is a practical starting policy for the production site:

data.json
[
  {
    "AllowedOrigins": ["https://viewparquet.com"],
    "AllowedMethods": ["GET", "HEAD"],
    "AllowedHeaders": ["*"],
    "ExposeHeaders": [
      "Accept-Ranges",
      "Content-Length",
      "Content-Range",
      "ETag"
    ],
    "MaxAgeSeconds": 3600
  }
]

Provider consoles use different syntax. Apply the equivalent policy on R2, GCS, or MinIO, then narrow AllowedHeaders to the Range and signing headers visible in your browser's OPTIONS preflight if your security policy requires it. Keep the origin specific instead of copying a wildcard-origin example into a sensitive bucket.

CORS controls which responses browser JavaScript can read; it does not grant object permission. IAM, bucket policy, HMAC/S3 credentials, and presigned authorization remain separate.

Verify range requests instead of assuming them

Open DevTools → Network, clear the list, then load or query the object:

  1. Confirm the object request targets your configured storage host.
  2. Inspect a GET request that includes a Range request header.
  3. A 206 Partial Content response with Content-Range demonstrates that the server honored that range.
  4. A 200 OK response to a ranged GET can mean the host ignored the range and returned the full object.
  5. Compare the transferred bytes while running a narrow projection versus a full scan.

When the endpoint supports HTTP range requests, DuckDB-WASM can read the footer and relevant row groups selectively when the Parquet statistics, row-group layout, and query allow it. Actual transfer depends on the query, selected columns, file layout, codecs, and endpoint behavior. A SELECT *, an export of every row, an unselective predicate, unsuitable row groups, missing useful statistics, or a host without range support can still transfer most or all of the file.

Troubleshooting by observable symptom

SymptomLikely causeWhat to check
Failed OPTIONS, “blocked by CORS,” or no readable responseCORS policyExact page origin, GET and HEAD, requested headers, exposed range headers, and whether the rule was applied to the correct bucket
403 AccessDenied or SignatureDoesNotMatchPermission or signing mismatchKey scope, s3:GetObject, region, endpoint, path/vhost style, and system/credential expiry
403 ExpiredTokenTemporary credentials expiredGenerate a new access key, secret, and STS session-token set
Presigned URL worked earlier but now returns 403URL or signing credentials expired; URL changedGenerate a fresh URL and paste the complete unmodified query string
404 NoSuchKey or NoSuchBucketWrong path or endpointBucket spelling, case-sensitive key, URL encoding, account endpoint, and actual file extension
416 Range Not SatisfiableHost or proxy mishandled the rangeProxy/CDN range support, object size/metadata, and whether a direct object-store URL behaves differently
Range request returns 200Host ignored the byte rangeExpect a larger transfer, fix the proxy/host, or download and open the file locally
R2 or GCS signing failureInteroperability settings mismatchEndpoint without https://, path-style addressing, region auto, and the correct S3/HMAC credential type
MinIO works locally but not on viewparquet.comMixed content, TLS, DNS, or CORSHTTPS endpoint, trusted certificate, public reachability, and production-origin CORS
File opens as the wrong formatAuto-detect had no usable extensionChoose Parquet explicitly; presigned URLs and extensionless paths can obscure the suffix
Query transfers far more than expectedFull scan or non-selective file layoutSelect only needed columns, add a useful filter, inspect row groups/statistics, and verify 206 responses

A browser cannot repair bucket CORS, issue missing IAM permission, renew an expired token, make an HTTP endpoint secure, or force a proxy to honor ranges. When one of those is the cause, change the storage configuration or download the object and open it as a local file.

After the file opens

The workbench uses the same table flow as a local source:

  • Run DuckDB SQL features supported by the shipped browser build.
  • Browse a paged results grid and the loaded-table schema snapshot.
  • Export supported results in the browser.
  • Use a narrow SELECT first; exporting all selected values may require reading them all.

Browser, extension, filesystem, endpoint, and codec support can differ from native DuckDB. For an automated or repeated workflow, test the equivalent query in the DuckDB CLI and keep the browser viewer for inspection.

Compare with DuckDB CLI

The local CLI model is similar:

query.sql
CREATE SECRET my_s3 (
  TYPE S3,
  PROVIDER config,
  KEY_ID '...',
  SECRET '...',
  REGION 'us-east-1'
);
 
SELECT count(*)
FROM read_parquet('s3://bucket/path/file.parquet');

In viewparquet, the dialog creates a scoped in-browser secret instead of writing credentials to a .duckdbrc file.

Primary references

Try it

Start with the controlled public sample, then connect your bucket once that path works. See the S3 Parquet viewer for the concise task page, the Parquet FAQ for answer-first troubleshooting, or Support when the browser error and storage logs still disagree.