Showing 36 of 361 projects
A complete, high-performance JSONPath implementation for Swift, enabling efficient querying and modification of JSON data.
A fluent builder for lazy streams and generators in Groovy, enabling functional-style data processing.
Feature generation code for the Kaggle Acquire Valued Shoppers Challenge, focusing on customer behavior prediction.
A unit test framework for Hive scripts that provides an embedded Hive environment with Derby database and HiveThriftService.
A Go library that converts XML documents into map[string]interface{} structures for flexible data handling.
A command-line JSON formatter and query tool with syntax highlighting and colorization.
Code samples demonstrating how to use popular applications on Amazon Elastic MapReduce (EMR).
A pure PHP library for reading and writing Parquet columnar storage files without external dependencies.
A fast, header-only CSV parser for modern C++ with a versatile API and support for custom types and error handling.
A Common Lisp library providing bit vector arithmetic, type conversions, and measurement functions.
A server application for managing GTFS transit data as part of the TRANSIT-Data-Tools suite.
A generalized map-reduce library for parallel processing on multicore systems, handling large files and infinite streams.
A machine learning solution for predicting job salaries from advertisements, developed for a Kaggle competition.
A Python SDK for deploying teams of AI research agents to forecast, score, classify, and gather data at scale.
A Go library and CLI tool for reading, writing, and processing transit data in GTFS and related formats.
A Go library and CLI tool for reading, writing, and processing transit data in GTFS and related formats.
An R package for processing and analyzing high-dimensional morphological profiling data from image-based cell biology.
Python tools to download, process, and analyze data from the Online Encyclopedia of Integer Sequences (OEIS).
A simple XML parser for Elixir designed to parse RSS/Atom feeds.
A Node.js framework for building machine-learning bots with a modular event handling and data processing pipeline.
A curated list of software, tools, and resources for exploring, organizing, and analyzing neuroimaging data with a focus on MRI.
A Python tool for stitching large volumetric images from light-sheet fluorescence microscopy.
A Node.js utility library for reading and processing GTFS public transit datasets with streaming and memory-efficient operations.
A Python parser for React's GraphQL query language, converting GraphQL syntax into structured Python dictionaries.
A Python tutorial demonstrating how to access and process Common Crawl's web archive datasets (WARC, WET, WAT) using tools like warcio, cdx_toolkit, and DuckDB.
A Common Lisp library providing functions and macros for manipulating arrays and performing numerical calculations.
A Python library providing tools to process and analyze OMOP-standardized clinical data from AP-HP's Clinical Data Warehouse.
A UNIX command line utility for parsing and formatting CSV files using printf-style syntax.
A framework for keeping biomedical text mining tools running on the latest publications from PubMed.
Generates GTFS shapes.txt and GeoJSON files from stop_times.txt using routing APIs like OSRM and Google Maps Directions.
An ActionScript 3.0 library for reading .xlsx Open XML Excel and Open Office spreadsheet files in Flash, Flex, and Air applications.
A reactive stream programming framework for live data computation with concurrent asynchronous operations.
An open-source Policy As Code Engine that programmatically creates and applies data policies to platforms like Snowflake, Databricks, and BigQuery.
A Meteor package that adds proper MongoDB aggregation support to Mongo.Collection instances.
A zero-dependency non-blocking buffered FIFO-pipeline library for Go that executes steps concurrently while maintaining output order.
A high-performance, streaming JSON parser for C applications requiring maximum speed with large JSON documents.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.