Data upload help

How to export and upload your data

Each service needs a specific export from your cloud. Find your service below, follow the steps for each item that applies to the clouds you use, and upload the files with the secure links we send you. Almost everything runs in your browser, in AWS CloudShell, Google Cloud Shell or a Databricks notebook, with nothing to install.

Find your service

Each card lists what that service needs, for the clouds and platforms you use, and links to the steps. Required: the service cannot start without it. Export, or grant access: required, but granting read access replaces the export. Optional: improves the result.

Before you start

  • After kickoff we email you one secure upload link per file, from info@finitizer.com, together with a ready-to-run command. Each link works for 7 days, accepts one file, and can only upload: nobody can read or list anything through it. If an email claiming to be from us asks for a password or a key, do not reply; we never ask for either.
  • Zip anything that is more than one file. A Cost and Usage Report, for example, is many files; send it as one zip.
  • AWS CloudShell (the >_ icon in the AWS console) and Google Cloud Shell (the >_ icon in the Google Cloud console) run in your browser and already have the AWS CLI or gcloud, zip and curl. CloudShell has about 1 GB of space and Cloud Shell about 5 GB; for larger exports use a computer with the same tools.
  • Words in CAPITALS in the commands, such as YOUR-BUCKET or region-YOUR-REGION, are placeholders: replace them with your own values before running.
  • No billing history yet? If you have to turn an export on today, it only holds data from today. Tell us the date: we agree with you whether to start on what there is or wait until enough months have built up.

AWS: Cost and Usage Report (CUR)

The line-by-line record of everything your AWS organization was billed, with the discounts and commitments applied. Run these steps from the management (payer) account.

About 15 minutes if you already have a report.

  1. 1

    Check whether you already have a report

    In the AWS console, open Billing and Cost Management, then Data Exports. If you see an export of type "Standard data export" (CUR 2.0) or "Legacy CUR export", note its S3 bucket, its path prefix and its name, and skip to step 3.

  2. 2

    If you have none, create one

    In Data Exports choose Create. Pick "Standard data export", give it a name, choose the CUR 2.0 table, tick "Include resource IDs", set time granularity to Hourly, keep Parquet, set file versioning to "Overwrite existing data export file", choose an S3 bucket you own, and create it. AWS starts delivering within 24 hours, from the current month only. You can ask AWS Support whether earlier months can be backfilled; tell us the date either way.

  3. 3

    Download the last three full months

    Open AWS CloudShell and run the commands below with your bucket, prefix and export name. CUR 2.0 keeps each month in a folder named BILLING_PERIOD=YYYY-MM; the first command lists them. Older reports use folders such as 20260701-20260801 or year=2026/month=7: sync those folders the same way. If your export keeps several deliveries per month (versioning "Create new data export file"), that is fine; tell us and we use only the latest one. For a portfolio review or a pre-acquisition review, list up to 12 months.

    AWS CloudShell
    aws s3 ls s3://YOUR-BUCKET/YOUR-PREFIX/YOUR-EXPORT-NAME/data/
    
    mkdir -p aws-cur
    for m in 2026-07 2026-08 2026-09; do
      aws s3 sync "s3://YOUR-BUCKET/YOUR-PREFIX/YOUR-EXPORT-NAME/data/BILLING_PERIOD=$m/" "aws-cur/BILLING_PERIOD=$m/"
    done
    du -sh aws-cur
  4. 4

    Zip it

    One file, ready for the upload link.

    AWS CloudShell
    zip -r aws-cur.zip aws-cur
  5. 5

    Upload it

    Run the upload command from our email in the same folder. See "Uploading a file" below.

Google Cloud: detailed billing export

Cloud Billing writes every charge, with credits and committed use discounts, to a BigQuery table. We need the detailed usage cost table, not the standard one.

About 15 minutes if the export is already on.

  1. 1

    Check that the detailed export is on

    In the Google Cloud console open Billing, choose your billing account, then Billing export and the BigQuery export tab. "Detailed usage cost" should say Enabled, with a project and a dataset. Note both. The table inside that dataset is named gcp_billing_export_resource_v1_ followed by your billing account ID.

  2. 2

    If it is off, turn it on

    Under "Detailed usage cost" choose Edit settings, pick a project and a dataset (create one called billing_export with location US if you have none) and save. Data starts arriving within a day, and Google fills in only a little past data, so turn it on as early as you can and tell us the date.

  3. 3

    Make sure you have a bucket

    The export is written to a Cloud Storage bucket in the same location as the dataset (for example US). If you do not have one, create it in Cloud Shell.

    Cloud Shell (only if you need a bucket)
    gcloud storage buckets create gs://YOUR-BUCKET --location=US --uniform-bucket-level-access
  4. 4

    Run the export query

    In BigQuery, open a new query tab, paste this, replace the names, and run it. It exports the last three full invoice months, so the totals match your invoices; change the 3 to 12 for a portfolio review or to 1 for training. For the BigQuery services, add a last line, AND service.description LIKE 'BigQuery%', to send only the BigQuery lines. The query reads the billing table once; for most companies that costs less than a dollar.

    BigQuery console
    EXPORT DATA OPTIONS (
      uri = 'gs://YOUR-BUCKET/finitizer/billing/billing-*.parquet',
      format = 'PARQUET', overwrite = true) AS
    SELECT *
    FROM `YOUR-PROJECT.YOUR-DATASET.gcp_billing_export_resource_v1_XXXXXX_XXXXXX_XXXXXX`
    WHERE invoice.month IN UNNEST(ARRAY(
      SELECT FORMAT_DATE('%Y%m', DATE_SUB(DATE_TRUNC(CURRENT_DATE(), MONTH), INTERVAL m MONTH))
      FROM UNNEST(GENERATE_ARRAY(1, 3)) AS m));
  5. 5

    Download, zip and tidy up

    Back in Cloud Shell. The last command deletes the export from your bucket once the zip is made.

    Cloud Shell
    gcloud storage cp -r gs://YOUR-BUCKET/finitizer/billing ./gcp-billing
    zip -r gcp-billing.zip gcp-billing
    gcloud storage rm -r gs://YOUR-BUCKET/finitizer/billing
  6. 6

    Upload it

    Run the upload command from our email in Cloud Shell. See "Uploading a file" below.

Another way: Prefer not to export at all? Grant our reader identity, which we send you at kickoff, read access to the billing dataset only: in BigQuery, open the dataset, choose Sharing, then Permissions, Add principal, paste the identity, choose the role BigQuery Data Viewer and save. We then read it in place, now and every month, and nothing has to be uploaded.

BigQuery: job, storage and slot history

Metadata about your queries and tables, read from INFORMATION_SCHEMA. No table data is included. Only needed if you export instead of granting us read access.

About 20 minutes.

  1. 1

    Note your projects and region

    List the projects in scope and the region your datasets live in, written the way BigQuery writes it: region-us, region-eu or region-us-central1, for example. Use it wherever the queries say region-YOUR-REGION. You need a Cloud Storage bucket in the same location (see the billing export steps above for how to create one).

  2. 2

    Export 30 days of jobs for each project

    Run this once per project in the BigQuery console. It includes the query text and who ran each job, which we use only to attribute cost and to rewrite your most expensive queries. To leave the text out, replace the last column, query, with TO_HEX(SHA256(query)) AS query_hash; we can still group repeated queries but cannot rewrite them. For BigQuery Slot Management, change INTERVAL 30 DAY to INTERVAL 90 DAY.

    BigQuery console, once per project
    EXPORT DATA OPTIONS (
      uri = 'gs://YOUR-BUCKET/finitizer/bq/jobs-YOUR-PROJECT-*.parquet',
      format = 'PARQUET', overwrite = true) AS
    SELECT creation_time, start_time, end_time, project_id, user_email, job_id, job_type,
           statement_type, priority, state, reservation_id, total_bytes_processed,
           total_bytes_billed, total_slot_ms, cache_hit, labels, query
    FROM `YOUR-PROJECT`.`region-YOUR-REGION`.INFORMATION_SCHEMA.JOBS_BY_PROJECT
    WHERE creation_time >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 30 DAY);
  3. 3

    Export storage and table layout for each project

    Two more statements per project: storage by table (for cold tables and the storage billing model), and which columns tables are partitioned or clustered on.

    BigQuery console, once per project
    EXPORT DATA OPTIONS (
      uri = 'gs://YOUR-BUCKET/finitizer/bq/storage-YOUR-PROJECT-*.parquet',
      format = 'PARQUET', overwrite = true) AS
    SELECT * FROM `YOUR-PROJECT`.`region-YOUR-REGION`.INFORMATION_SCHEMA.TABLE_STORAGE;
    
    EXPORT DATA OPTIONS (
      uri = 'gs://YOUR-BUCKET/finitizer/bq/layout-YOUR-PROJECT-*.parquet',
      format = 'PARQUET', overwrite = true) AS
    SELECT table_catalog, table_schema, table_name, column_name,
           is_partitioning_column, clustering_ordinal_position
    FROM `YOUR-PROJECT`.`region-YOUR-REGION`.INFORMATION_SCHEMA.COLUMNS
    WHERE is_partitioning_column = 'YES' OR clustering_ordinal_position IS NOT NULL;
  4. 4

    For BigQuery Slot Management: per-minute slot usage

    Once per project, 90 days of slot usage summed by minute, which is what reservations are sized from. Skip this for the cost audit.

    BigQuery console, once per project
    EXPORT DATA OPTIONS (
      uri = 'gs://YOUR-BUCKET/finitizer/bq/slots-YOUR-PROJECT-*.parquet',
      format = 'PARQUET', overwrite = true) AS
    SELECT TIMESTAMP_TRUNC(period_start, MINUTE) AS minute, project_id, reservation_id,
           job_type, SUM(period_slot_ms) / 60000 AS avg_slots, COUNT(DISTINCT job_id) AS jobs
    FROM `YOUR-PROJECT`.`region-YOUR-REGION`.INFORMATION_SCHEMA.JOBS_TIMELINE_BY_PROJECT
    WHERE job_creation_time >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 90 DAY)
    GROUP BY minute, project_id, reservation_id, job_type;
  5. 5

    Download, zip and upload

    In Cloud Shell, then run the upload command from our email.

    Cloud Shell
    gcloud storage cp -r gs://YOUR-BUCKET/finitizer/bq ./bq-history
    zip -r bq-history.zip bq-history
    gcloud storage rm -r gs://YOUR-BUCKET/finitizer/bq

Another way: Quicker: grant our reader identity BigQuery Resource Viewer and BigQuery Metadata Viewer on each project in scope (and, if you use editions, BigQuery Resource Viewer on the reservation administration project). These roles read; they cannot change anything. We then collect all of this ourselves and nothing has to be uploaded.

BigQuery: reservations, assignments and commitments

Your current capacity setup, read from the reservation administration project: the project where your reservations and commitments live. Only if you use BigQuery editions, and only if you export instead of granting access.

About 10 minutes.

  1. 1

    Export the four views

    Run this in the BigQuery console in the administration project, with your project name and region. If a statement returns no rows (you hold no commitments, for example), that is fine.

    BigQuery console, administration project
    EXPORT DATA OPTIONS (
      uri = 'gs://YOUR-BUCKET/finitizer/res/timeline-*.parquet', format = 'PARQUET', overwrite = true) AS
    SELECT * FROM `ADMIN-PROJECT`.`region-YOUR-REGION`.INFORMATION_SCHEMA.RESERVATIONS_TIMELINE
    WHERE period_start >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 90 DAY);
    
    EXPORT DATA OPTIONS (
      uri = 'gs://YOUR-BUCKET/finitizer/res/reservations-*.parquet', format = 'PARQUET', overwrite = true) AS
    SELECT * FROM `ADMIN-PROJECT`.`region-YOUR-REGION`.INFORMATION_SCHEMA.RESERVATIONS;
    
    EXPORT DATA OPTIONS (
      uri = 'gs://YOUR-BUCKET/finitizer/res/assignments-*.parquet', format = 'PARQUET', overwrite = true) AS
    SELECT * FROM `ADMIN-PROJECT`.`region-YOUR-REGION`.INFORMATION_SCHEMA.ASSIGNMENTS;
    
    EXPORT DATA OPTIONS (
      uri = 'gs://YOUR-BUCKET/finitizer/res/commitments-*.parquet', format = 'PARQUET', overwrite = true) AS
    SELECT * FROM `ADMIN-PROJECT`.`region-YOUR-REGION`.INFORMATION_SCHEMA.CAPACITY_COMMITMENTS;
  2. 2

    Download, zip and upload

    In Cloud Shell, then run the upload command from our email.

    Cloud Shell
    gcloud storage cp -r gs://YOUR-BUCKET/finitizer/res ./bq-reservations
    zip -r bq-reservations.zip bq-reservations
    gcloud storage rm -r gs://YOUR-BUCKET/finitizer/res

Another way: Quicker: grant our reader identity BigQuery Resource Viewer on the administration project. It reads reservations and commitments and cannot change or buy anything.

Databricks: system tables

Billing, compute, query and job history from Unity Catalog system tables. Needs Unity Catalog with system tables enabled. Only needed if you export instead of creating a service principal for us.

About 20 minutes.

  1. 1

    Create a volume to export into

    In a Databricks SQL editor or notebook, with a catalog and schema you can write to.

    Databricks SQL
    CREATE VOLUME IF NOT EXISTS YOUR_CATALOG.YOUR_SCHEMA.finitizer_export;
  2. 2

    Export 30 days of system tables as one zip

    Paste this into a Python notebook cell, set the catalog and schema on the first line, and run it. It writes each table as Parquet, then zips them into one file in the volume. A table your workspace does not have is skipped and reported. To leave out query text, change the query history line to SELECT * EXCEPT (statement_text).

    Databricks notebook (Python)
    import shutil
    out = "/Volumes/YOUR_CATALOG/YOUR_SCHEMA/finitizer_export"
    queries = {
        "billing_usage": "SELECT * FROM system.billing.usage WHERE usage_date >= date_sub(current_date(), 30)",
        "list_prices": "SELECT * FROM system.billing.list_prices",
        "clusters": "SELECT * FROM system.compute.clusters",
        "warehouses": "SELECT * FROM system.compute.warehouses",
        "query_history": "SELECT * FROM system.query.history WHERE start_time >= current_timestamp() - INTERVAL 30 DAYS",
        "jobs": "SELECT * FROM system.lakeflow.jobs",
        "job_runs": "SELECT * FROM system.lakeflow.job_run_timeline WHERE period_start_time >= current_timestamp() - INTERVAL 30 DAYS",
    }
    for name, sql in queries.items():
        try:
            spark.sql(sql).write.mode("overwrite").parquet(f"{out}/tables/{name}")
            print("exported", name)
        except Exception as e:
            print("skipped", name, str(e)[:120])
    # Zip on local disk, then copy the single file into the volume.
    shutil.make_archive("/tmp/databricks-export", "zip", f"{out}/tables")
    shutil.copy("/tmp/databricks-export.zip", f"{out}/databricks-export.zip")
    print("ready:", f"{out}/databricks-export.zip")
  3. 3

    Download the zip

    In Catalog Explorer, open your catalog, schema and the finitizer_export volume, select databricks-export.zip and choose Download. It is one file, so it downloads in one click.

  4. 4

    Upload it

    From the folder you downloaded it to, run the upload command from our email (curl is built into Windows 10 and later, macOS and Linux). See "Uploading a file" below. Then delete the finitizer_export volume.

Another way: Quicker: create a service principal with SELECT on the system schemas (billing, compute, query, lakeflow) and permission to use one SQL warehouse, and share its token through the upload link. Revoke the token when the audit ends.

Private equity portfolio: what each company sends

Each portfolio company sends its own data, through its own upload links, so no company sees another's. The fund only needs to tell us who to contact at each company and the exit multiple to use.

About 2 hours per company.

  1. 1

    Billing data, 3 to 12 months

    Best: the AWS Cost and Usage Report or the Google Cloud detailed billing export, produced with the steps above (list up to 12 months in the AWS loop; change the 3 to 12 in the Google Cloud query). A company without an export can send 12 monthly cloud invoices and a monthly cost report by service and by account instead (AWS Cost Explorer grouped by Service and then by Linked account, monthly, as CSV; or the Google Cloud Billing reports page grouped by service and then by project, as CSV).

  2. 2

    Revenue and EBITDA

    Each company's monthly revenue and EBITDA for the same period, so we can show cloud cost as a share of revenue and the effect of savings on EBITDA. If the fund prefers to supply these figures, the company can skip this step.

  3. 3

    Contracts, if any

    Any enterprise discount or private pricing agreement, Google Cloud commit, Marketplace contract, or Snowflake or Databricks contract, so savings are priced at the rates the company actually pays.

  4. 4

    For a Company Value Plan: read-only access

    The company deploys our read-only role with one of our engineers on a short call, the same setup as the Cloud Cost Audit. Nothing to do for the Portfolio Estimate.

  5. 5

    Buying a new company?

    Before close, ask the target for the same billing data and contracts in the data room, and give us access to that folder or upload the files with our links. Read-only access waits until signing allows it.

Training: a sample of your billing data

The labs run on your own numbers. One recent month is enough, and you can mask it.

About 15 minutes.

  1. 1

    Export one month

    Follow the AWS or Google Cloud steps above for a single month: on AWS sync only the most recent BILLING_PERIOD folder; on Google Cloud change the 3 in the query to 1.

  2. 2

    For the AI and LLM module

    A recent month of usage from your AI provider: most provider consoles (Anthropic, OpenAI, Amazon Bedrock, Vertex AI) offer a usage or cost download as CSV on their usage or billing page. The AI lines of your cloud bill also work.

  3. 3

    Mask it if you like

    Account names, project names and tags can be replaced with placeholders; the labs need the shape of your spend, not who owns it. If you cannot share any billing data, tell us: the labs then use a realistic sample estate.

  4. 4

    Upload it

    Zip it and run the upload command from our email.

Discount, private pricing and credit terms

Optional, but it means savings are priced at the rates you actually pay rather than list prices.

A few minutes.

  1. 1

    Send the documents

    The summary or rate schedule of any AWS Enterprise Discount Program or private pricing agreement, Google Cloud discount or commit agreement, and active credits. A PDF is fine. Use a separate upload link from the billing data; ask us for one, or attach it to a reply to our email if you prefer.

Uploading a file

Each upload link from our email comes with a ready-to-run command. Run it in the folder that holds the file.

AWS CloudShell, Google Cloud Shell, macOS or Linux

Terminal
curl -X PUT -H "Content-Type: application/octet-stream" --upload-file aws-cur.zip "THE-LINK-FROM-OUR-EMAIL"

Windows (PowerShell): note curl.exe

PowerShell
curl.exe -X PUT -H "Content-Type: application/octet-stream" --upload-file aws-cur.zip "THE-LINK-FROM-OUR-EMAIL"

When it finishes, curl prints a short reply that includes "size": followed by the number of bytes. That means the file arrived. Reply to our email to let us know.

If something goes wrong

The reply says the upload ID is invalid, or 404
The link expired (after 7 days) or was already used. Ask us for a new one.
The upload stopped part way
Ask us for a new link and run it again; the partial file is discarded.
Windows says a parameter cannot be found
You ran curl instead of curl.exe in PowerShell. Use curl.exe, exactly as in the Windows command.
CloudShell runs out of space
AWS CloudShell has about 1 GB. Run the same commands on a computer with the AWS CLI, or export fewer months and send them in two files.
An export query fails with a location error
The bucket must be in the same location as the data. Create a bucket with the same location as the dataset, or change region-YOUR-REGION to the region your datasets use.
Your laptop blocks command-line tools
Run everything in AWS CloudShell or Google Cloud Shell, which run in the browser, or book a short call and we will walk through it with you.

How your data is protected

  • Every customer gets a separate storage bucket in Finitizer's Google Cloud environment in the US, encrypted at rest and in transit, with public access blocked.
  • Upload links can only write. Nobody, including someone who intercepts a link, can read, list or download anything through it.
  • Files are used only for your engagement, deleted automatically 90 days after upload, and deleted sooner on request. Deletion is immediate and final.
  • We sign your NDA before any data is shared. Query text and user names in BigQuery or Databricks history are used only to attribute cost.

Stuck on a step?

Email info@finitizer.com with the step and what you see, or book a short call and one of our FinOps Engineers will walk through it with you.

Book a free 30-minute call