Showing 36 of 361 projects
Provides HBase adapters for reading and writing data within Cascading data processing workflows on Hadoop clusters.
An open-source toolkit for analyzing line-oriented JSON Twitter archives using Apache Spark.
A thin and fast C library for parsing static GTFS (General Transit Feed Specification) feeds.
A tool that merges multiple GTFS transit datasets into a single unified feed for regional transit analysis.
A collection of scripts for processing Merck molecular activity challenge data from Kaggle into formats suitable for machine learning.
An Elixir package for low-memory consumption iteration over large BSON files by applying a function to each document.
A Swift library providing a generic table view controller with external data processing for iOS apps.
A Go package for converting entire slices between primitive types and their string representations.
Generates a static R-tree spatial index from GeoJSON features for fast geographic queries.
A high-performance LTSV (Labeled Tab-Separated Value) parser library for Go.
A fast, memory-efficient CLI tool for splitting large SQL dump files into individual table files and converting between SQL dialects.
A column-oriented data processing engine that replaces joins and group-by with functional definitions for batch and stream analytics.
Automates AWS EMR cluster deployment and Spark job submission with integrated logging and PySparkling support for H2O.
A multi-extract, multi-level dataset of Mozilla Bugzilla issue tracking history spanning 15 years for software engineering research.
Angular service for saving data to CSV files with customizable mapping and formatting.
Command line tool to extract shapes from a GTFS dataset.
A Java API for the Onyx Platform, providing Java equivalents for workflows, utilities for Clojure maps, and tools for core.async plugins.
A CSV parsing and printing library for ActionScript, ported from Apache Commons CSV with incremental parsing support.
A tool that computes transit service areas from static GTFS data and outputs them as GeoJSON files.
Flow analysis on disaggregated bilateral trade data using Golem's decentralized computing network.
A Vert.x facade for Jolt that enables JSON-to-JSON transformations using Vert.x JsonObject instead of Map<String,Object>.
A library for performing mathematical operations on number arrays stored in binaries, with support for handling missing values.
A Swift library providing a generic UICollectionViewController with external data processing and flexible cell configuration.
A fast, low-memory .NET library for importing, exporting, and templating Excel spreadsheets with streaming row-by-row processing.
A promise-based wrapper for through2 that enables asynchronous mapping over Node.js streams.
A JavaScript library that merges single-type GeoJSON features into a multi-type GeoJSON feature with customizable property aggregation.
Generates yo mama jokes using Markov chains trained on a collected dataset of jokes.
A Node.js library for parsing CDXJ files produced by web archiving tools like Pywb.
A thin C# gRPC client for communicating with Apache Spark Connect servers, enabling .NET applications to interact with Spark clusters.
Streaming Node.js tool that adds unique IDs to GeoJSON features, enabling large file processing.
A sample Spark job demonstrating how to use Spark Jobserver to run Apache Spark analytics with Cassandra.
A sample Spark job demonstrating analytics on Cassandra with SSL encryption for secure data processing.
Read and write OpenDocument Spreadsheet (ODS) files in R, supporting both ordinary and Flat ODS formats.
A V library for reading and parsing Excel XLSX files.
A high-performance Erlang NIF CSV parser and writer based on libcsv for processing large data volumes.
A Java project that reads from and writes to TDengine using Apache Spark for data processing.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.