~/tools/trino
Trino
tool

Trino

Trino is an open source distributed SQL query engine you can self-host — a query-federation layer that runs fast ANSI SQL across data lakes, warehouses, and databases at once, standing in for services like Amazon Athena, Snowflake, and BigQuery.

What is Trino?

Trino is a distributed SQL query engine that runs fast, interactive analytics across many data sources at once — without moving or copying the data. It stores nothing itself; instead it connects to object storage, warehouses, and databases through pluggable connectors and queries them in place using standard ANSI SQL. Trino began as PrestoSQL, forked from Facebook’s Presto and rebranded in 2020.

What is Trino best for?

Trino is best for interactive, ad-hoc analytics over large data lakes and for federated queries that join data spread across separate systems. Teams point BI tools like Tableau or Power BI at it, run exploratory SQL over data in S3 or HDFS, and use it as the query layer for lakehouse tables. It suits data platform and analytics engineering teams, and it scales from a laptop to exabyte-scale clusters used by Netflix, LinkedIn, Stripe, and Shopify.

What can Trino do?

  • Query data in place across 50+ connectors — Hive, Iceberg, Delta Lake, S3, PostgreSQL, MySQL, Cassandra, MongoDB, Kafka, and more.
  • Federate queries — join a table in PostgreSQL against files in S3 and a collection in MongoDB in a single SQL statement.
  • Run ANSI SQL with window functions, CTEs, and complex joins, so existing SQL and BI tools work unchanged.
  • Separate compute from storage, scaling a coordinator-and-workers cluster independently of where the data lives.
  • Serve mixed workloads — sub-second dashboards, interactive exploration, and long-running batch ETL from the same engine.
  • Connect standard tools over JDBC/ODBC and integrate with lakehouse table formats and catalogs.

Where does Trino fall short?

  • It’s not a database. Trino has no storage, no durable tables of its own, and limited support for updates, deletes, and transactions — it’s a read-heavy query layer over other systems, not a system of record.
  • Memory-bound execution. Trino processes queries in memory and can fail large joins or aggregations that exceed cluster RAM, so heavy workloads need careful tuning and generous memory rather than the near-infinite elasticity of a serverless warehouse.
  • Modern runtime requirement. Recent releases require a current JDK (Java 25 for the latest versions), which can complicate deployment on environments standardized around older Java.

Is Trino free?

Yes — Trino is fully free and open source under the Apache 2.0 license, with no paid tier or enterprise edition from the project itself. You run it on your own hardware or cloud at no license cost. Commercial support and managed Trino services are sold separately by third parties such as Starburst, but the engine is free to self-host.

What does Trino replace?

Trino stands in for serverless and cloud analytics query services when you’d rather own the engine and query your own storage. It’s a self-hosted alternative to Amazon Athena (which is itself built on Presto/Trino), and it overlaps with the interactive-query role of Snowflake, Google BigQuery, and Databricks — though those bundle managed storage and infrastructure that Trino leaves to you.

FAQ

Is Trino open source? Yes. Trino is licensed under Apache 2.0 and governed by the non-profit Trino Software Foundation. The full source is on GitHub and free to use, modify, and self-host.

Is Trino the same as Presto? Trino is the renamed continuation of PrestoSQL, forked from the original Presto project and rebranded to Trino in 2020. It’s maintained by Presto’s original creators and has diverged from Meta’s Presto since.

Can I self-host Trino for free? Yes. Trino is free to run yourself — you deploy a coordinator and worker nodes on your own servers or cloud and pay only for that infrastructure. There’s no license fee; only optional third-party support costs money.

What do I need to run Trino? A cluster of one coordinator plus one or more worker nodes running a 64-bit JDK (Java 25 for current releases), on Linux or macOS. A single node works for testing, but production analytics needs multiple workers and enough RAM to hold query working sets in memory.