From bbdde3d8dd774870a9e0b9723f05bb6da7b34b35 Mon Sep 17 00:00:00 2001 From: Steve Nyemba Date: Mon, 7 Sep 2026 11:49:18 -0500 Subject: [PATCH] documentation --- README.md | 64 +++++++++++++++++++++++++++---------------------------- 1 file changed, 31 insertions(+), 33 deletions(-) diff --git a/README.md b/README.md index 4b0f214..d2089df 100644 --- a/README.md +++ b/README.md @@ -1,44 +1,42 @@ -# Introduction +# Data-Transport -This project implements an abstraction of objects that can have access to a variety of data stores, implementing read/write with a simple and expressive interface. This abstraction works with **NoSQL**, **SQL** and **Cloud** data stores and leverages **pandas**. +A powerful abstraction layer for seamless data communication across diverse systems. **data-transport** allows you to interact with NoSQL, SQL, and Cloud storage using a consistent interface powered by **Pandas** and **SQLAlchemy**. -# Why Use Data-Transport ? +## Why Choose data-transport? -Data transport is a simple framework that enables read/write to multiple databases or technologies that can hold data. In using **data-transport**, you are able to: +* **Unified Interface:** Connect to PostgreSQL, MySQL, MongoDB, S3, etc., using the same consistent code. +* **Security First:** Prevents the dissipation of database connectivity information to protect against security breaches. +* **Simplicity & Power:** Leverages Pandas DataFrames and SQLAlchemy for intuitive data manipulation. +* **Robust Pipelines:** Easily integrate pre-processing and post-processing as unified pipelines. +* **CLI Integration:** Includes a dedicated CLI for registry management and ETL task execution. -- Enjoy the simplicity of **data-transport** because it leverages SQLAlchemy & Pandas data-frames. -- Share notebooks and code without having to disclosing database credentials. -- Seamlessly and consistently access to multiple database technologies at no cost -- No need to worry about accidental writes to a database leading to inconsistent data -- Implement consistent pre and post processing as a pipeline i.e aggregation of functions -- **data-transport** is open-source under MIT License https://github.com/lnyemba/data-transport +## Supported Features -## Installation - -Within the virtual environment perform the following, the options for installation are: - -**sql** - by default postgresql, mysql, sqlserver, sqlite3+, duckdb - - pip install data-transport[cloud,nosql,other,all]git+https://github.com/lnyemba/data-transport.git - -Options to install components in square brackets, these components are +| Component | Technologies Covered | +| :--- | :--- | +| **SQL** | PostgreSQL, MySQL, SQL Server, SQLite3+, DuckDB | +| **NoSQL** | MongoDB, CouchDB | +| **Warehouse** | Apache Iceberg, Apache Drill | +| **Cloud** | Nextcloud, S3 | +| **Other** | Files, RabbitMQ, HTTP | -**warehouse** - Apache Iceberg, Apache Drill - -**cloud**  - to support nextcloud, s3 - -**nosql** - support for mongodb, couchdb - -**other**  - support for files, rabbitmq, http +## Installation - pip install data-transport[nosql,cloud,warehouse,all]@git+https://github.com/lnyemba/data-transport.git +Install the core package and your desired components: -## Additional features +```bash +# Basic installation with default SQL support +pip install data-transport@git+https://github.com/lnyemba/data-transport - - In addition to read/write, there is support for functions for pre/post processing - - CLI interface to add to registry, run ETL - - scales and integrates into shared environments like apache zeppelin; jupyterhub; SageMaker; ... +# Full suite (SQL, NoSQL, Cloud, Warehouse) +pip install "data-transport[nosql,cloud,warehouse,all]"@git+https://github.com/lnyemba/data-transport.git +``` -## Learn More +## Advanced Capabilities +* **Automated Pipelines:** Seamlessly aggregate functions for data cleaning and transformation. +* **Portability:** Share notebooks and scripts without exposing raw credentials. +* **Scalable Integration:** Compatible with environments like Apache Zeppelin, JupyterHub, and SageMaker. -We have available notebooks with sample code to read/write against mongodb, couchdb, Netezza, PostgreSQL, Google Bigquery, Databricks, Microsoft SQL Server, MySQL ... Visit [data-transport homepage](https://healthcareio.the-phi.com/data-transport) +--- +[Learn More at the Project Website](https://healthcareio.the-phi.com/data-transport) +License: [MIT](https://github.com/lnyemba/data-transport)