Showing 5 of 5 projects
A Python library that simplifies data integration between pandas and AWS services like Athena, S3, Redshift, and more.
A Python library that simplifies data integration between pandas and AWS services like Athena, S3, Redshift, and more.
An open specification for storing geospatial vector data (points, lines, polygons) in the Apache Parquet columnar storage format.
A genomics analysis platform that uses Apache Spark to parallelize genomic data processing across clusters, replacing traditional file-based workflows.
A massively-parallel C++ SQL query engine for lightning-fast analytics on petabytes of data in Hadoop clusters.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.