Showing 25 of 25 projects
An in-process analytical SQL database management system designed for high-performance data analysis.
A scalable time series database optimized for real-time metrics, events, and analytics with fast query response.
A scalable time series database optimized for real-time metrics, events, and analytics with fast query response.
An open-source storage framework that enables building a Lakehouse architecture with ACID transactions and scalable metadata handling.
A command-line tool for running SQL queries against JSON, CSV, Excel, Parquet, and other structured data files.
A blazing-fast command-line toolkit for querying, slicing, analyzing, transforming, and validating tabular data (CSV, Excel, JSONL, etc.).
A lightweight TUI application for viewing and querying tabular data files like CSV, Parquet, and JSON with SQL support.
A graph database framework for storing and querying large-scale graphs with rich properties and in-database aggregation.
An embedded database for serverless and edge runtimes, storing data as Parquet on S3 with stateless compute.
A genomics analysis platform that uses Apache Spark to parallelize genomic data processing across clusters, replacing traditional file-based workflows.
A simple, fast, and flexible ETL framework for .NET with built-in readers and writers for CSV, JSON, XML, Parquet, and more.
Global open dataset of aggregated fixed and mobile network performance metrics (download/upload/latency) in geospatial tiles.
A Go library that generates type-safe Parquet readers and writers from Go structs or existing Parquet files.
A Spark application for migrating data to ScyllaDB from CQL-compatible databases or DynamoDB via Alternator.
An easy-to-use Python feature store for machine learning, optimized for timeseries data and built on Dask.
A pure PHP library for reading and writing Parquet columnar storage files without external dependencies.
A next-generation data analysis library for Go, offering parallel processing, data visualization, and seamless Python integration as an alternative to Pandas.
Common Lisp CFFI wrapper around the DuckDB C API
A Python tutorial demonstrating how to access and process Common Crawl's web archive datasets (WARC, WET, WAT) using tools like warcio, cdx_toolkit, and DuckDB.
A public dataset of Ethereum network events including beacon chain, mempool, and canonical chain data for analysis.
Global flight schedule datasets extracted from ADS-B position transmissions, published quarterly from 2024 onwards.
A command-line tool that scrapes, normalizes, and archives real-time public transit (GTFS Realtime) data for historical analysis.
A VS Code extension for viewing large datasets (JSONL/Parquet/CSV) instantly without crashes, with 16 production LLM tokenizers for accurate token counting.
A blazingly fast data comparison tool for Python that instantly compares massive CSV/Parquet datasets, powered by Rust.
A Python package for converting PCAP network capture files to Parquet, CSV, or JSON formats.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.