|
|
|
@ -1,44 +1,42 @@
|
|
|
|
# Introduction
|
|
|
|
# Data-Transport
|
|
|
|
|
|
|
|
|
|
|
|
This project implements an abstraction of objects that can have access to a variety of data stores, implementing read/write with a simple and expressive interface. This abstraction works with **NoSQL**, **SQL** and **Cloud** data stores and leverages **pandas**.
|
|
|
|
A powerful abstraction layer for seamless data communication across diverse systems. **data-transport** allows you to interact with NoSQL, SQL, and Cloud storage using a consistent interface powered by **Pandas** and **SQLAlchemy**.
|
|
|
|
|
|
|
|
|
|
|
|
# Why Use Data-Transport ?
|
|
|
|
## Why Choose data-transport?
|
|
|
|
|
|
|
|
|
|
|
|
Data transport is a simple framework that enables read/write to multiple databases or technologies that can hold data. In using **data-transport**, you are able to:
|
|
|
|
* **Unified Interface:** Connect to PostgreSQL, MySQL, MongoDB, S3, etc., using the same consistent code.
|
|
|
|
|
|
|
|
* **Security First:** Prevents the dissipation of database connectivity information to protect against security breaches.
|
|
|
|
|
|
|
|
* **Simplicity & Power:** Leverages Pandas DataFrames and SQLAlchemy for intuitive data manipulation.
|
|
|
|
|
|
|
|
* **Robust Pipelines:** Easily integrate pre-processing and post-processing as unified pipelines.
|
|
|
|
|
|
|
|
* **CLI Integration:** Includes a dedicated CLI for registry management and ETL task execution.
|
|
|
|
|
|
|
|
|
|
|
|
- Enjoy the simplicity of **data-transport** because it leverages SQLAlchemy & Pandas data-frames.
|
|
|
|
## Supported Features
|
|
|
|
- Share notebooks and code without having to disclosing database credentials.
|
|
|
|
|
|
|
|
- Seamlessly and consistently access to multiple database technologies at no cost
|
|
|
|
|
|
|
|
- No need to worry about accidental writes to a database leading to inconsistent data
|
|
|
|
|
|
|
|
- Implement consistent pre and post processing as a pipeline i.e aggregation of functions
|
|
|
|
|
|
|
|
- **data-transport** is open-source under MIT License https://github.com/lnyemba/data-transport
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
## Installation
|
|
|
|
| Component | Technologies Covered |
|
|
|
|
|
|
|
|
| :--- | :--- |
|
|
|
|
Within the virtual environment perform the following, the options for installation are:
|
|
|
|
| **SQL** | PostgreSQL, MySQL, SQL Server, SQLite3+, DuckDB |
|
|
|
|
|
|
|
|
| **NoSQL** | MongoDB, CouchDB |
|
|
|
|
**sql** - by default postgresql, mysql, sqlserver, sqlite3+, duckdb
|
|
|
|
| **Warehouse** | Apache Iceberg, Apache Drill |
|
|
|
|
|
|
|
|
| **Cloud** | Nextcloud, S3 |
|
|
|
|
pip install data-transport[cloud,nosql,other,all]git+https://github.com/lnyemba/data-transport.git
|
|
|
|
| **Other** | Files, RabbitMQ, HTTP |
|
|
|
|
|
|
|
|
|
|
|
|
Options to install components in square brackets, these components are
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
**warehouse** - Apache Iceberg, Apache Drill
|
|
|
|
## Installation
|
|
|
|
|
|
|
|
|
|
|
|
**cloud** - to support nextcloud, s3
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
**nosql** - support for mongodb, couchdb
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
**other** - support for files, rabbitmq, http
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
pip install data-transport[nosql,cloud,warehouse,all]@git+https://github.com/lnyemba/data-transport.git
|
|
|
|
Install the core package and your desired components:
|
|
|
|
|
|
|
|
|
|
|
|
## Additional features
|
|
|
|
```bash
|
|
|
|
|
|
|
|
# Basic installation with default SQL support
|
|
|
|
|
|
|
|
pip install data-transport@git+https://github.com/lnyemba/data-transport
|
|
|
|
|
|
|
|
|
|
|
|
- In addition to read/write, there is support for functions for pre/post processing
|
|
|
|
# Full suite (SQL, NoSQL, Cloud, Warehouse)
|
|
|
|
- CLI interface to add to registry, run ETL
|
|
|
|
pip install "data-transport[nosql,cloud,warehouse,all]"@git+https://github.com/lnyemba/data-transport.git
|
|
|
|
- scales and integrates into shared environments like apache zeppelin; jupyterhub; SageMaker; ...
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
## Learn More
|
|
|
|
## Advanced Capabilities
|
|
|
|
|
|
|
|
* **Automated Pipelines:** Seamlessly aggregate functions for data cleaning and transformation.
|
|
|
|
|
|
|
|
* **Portability:** Share notebooks and scripts without exposing raw credentials.
|
|
|
|
|
|
|
|
* **Scalable Integration:** Compatible with environments like Apache Zeppelin, JupyterHub, and SageMaker.
|
|
|
|
|
|
|
|
|
|
|
|
We have available notebooks with sample code to read/write against mongodb, couchdb, Netezza, PostgreSQL, Google Bigquery, Databricks, Microsoft SQL Server, MySQL ... Visit [data-transport homepage](https://healthcareio.the-phi.com/data-transport)
|
|
|
|
---
|
|
|
|
|
|
|
|
[Learn More at the Project Website](https://healthcareio.the-phi.com/data-transport)
|
|
|
|
|
|
|
|
License: [MIT](https://github.com/lnyemba/data-transport)
|
|
|
|
|