Workspace

The Analysis Capabilities

Real Analysis, on a Real Scientific Stack.

Every capability below runs in an isolated analysis cloud computer with a scientific Python stack and durable project outputs.

16

Analysis Domains

67

Capabilities

35

Research Packages

How Analysis Runs Happen

A lab environment, not a code snippet box.

Each analysis command starts an isolated job, stages the project from private cloud storage, runs without network access, and synchronizes changed outputs back to the project.

Real Cloud Computer

A private microVM per user

Full Python, Node, and shell, not a snippet runner

Research Environment

Read-only, reproducible

Every project runs the same vetted scientific stack; nothing installs at runtime

Compute Envelope

Sufficient RAM

Sized for interactive research turns; heavy jobs chunk or stream

No Network Inside

Your data stays in the sandbox

Bulk collection runs host-side; analysis works on files by path

Provenance by Default

Raw data is immutable

Every derivation writes a new, clearly named file — rerunnable anytime

Saved as Artifacts

Scripts, figures, notebooks

Each run lands in your project workspace next to the data

Statistics & Inference

From first descriptives to models with defensible uncertainty.

Descriptive Statistics

Know your data before modeling it.

  • Summary Statistics

    Means, medians, spread, quartiles — missing values accounted

  • Frequency Tables

    Level counts and percentages, reconciled and sorted

  • Dataset Profiling

    Row/column census with types, missing, distinct, samples

  • Cross-Tabulation

    Contingency tables for categorical relationships

  • Distribution Shape

    Skewness, kurtosis, and quantile checks

Hypothesis Testing

Parametric, non-parametric, and everything the assumptions require.

  • t-tests & ANOVA

    Group comparisons with post-hoc follow-ups

  • Non-Parametric Tests

    Mann-Whitney, Wilcoxon, Kruskal-Wallis, χ², Kolmogorov-Smirnov

  • Correlation with CIs

    Pearson, Spearman, and Kendall with significance

  • Effect Sizes & Intervals

    Estimates with confidence, not just p-values

  • Power Analysis

    Sample sizing and minimum detectable effects

  • Multiple-Testing Control

    Bonferroni, Holm, and FDR corrections

Regression & Generalized Models

Models with full inference — coefficients, intervals, and diagnostics.

  • Linear Models

    OLS/WLS/GLS with R², F-tests, residual diagnostics

  • Generalized Linear Models

    Logistic, Poisson, and negative binomial

  • Robust Standard Errors

    HC estimators and cluster-robust inference

  • Mixed-Effects Models

    Random intercepts and slopes for multilevel data

  • GEE

    Correlated and longitudinal designs

  • Nonlinear Curve Fitting

    Custom models fit with confidence bands

Time Series & Forecasting

Trends, seasonality, and forecasts from your own series.

  • ARIMA & SARIMAX

    Seasonal models with exogenous regressors

  • Exponential Smoothing

    ETS state-space forecasting

  • Decomposition

    Trend, seasonal, and remainder separation

  • Vector Autoregression

    Multi-series dynamics and impulse response

  • Signal Processing

    Filters, spectra, and peak detection

Survival & Duration Analysis

Time-to-event modeling that respects censoring.

  • Kaplan-Meier

    Survival curves with confidence bands

  • Cox Proportional Hazards

    Semiparametric regression with diagnostics

  • Parametric Survival

    Accelerated failure-time families

Machine Learning

Classical, tabular, CPU-scale, the kind of ML research questions actually need.

Predictive Modeling

Train, validate, and persist — the full supervised workflow.

  • Classification & Regression

    Logistic, SVM, k-NN, naive Bayes

  • Tree Ensembles

    Random forests and gradient boosting

  • Neural Baselines

    MLPs for small tabular problems

  • Cross-Validation

    K-fold, stratified, and grouped splits

  • Hyperparameter Search

    Grid and randomized tuning

  • Evaluation & Calibration

    ROC/AUC, PR curves, confusion matrices

  • Model Persistence

    Fitted models saved to your workspace

Unsupervised Learning

Find structure when nobody handed you labels.

  • Clustering

    k-means, DBSCAN, and agglomerative

  • Gaussian Mixtures

    EM-fitted soft clustering

  • Dimensionality Reduction

    PCA, t-SNE, and MDS

  • Anomaly Detection

    Isolation forests and density methods

  • Cluster Diagnostics

    Silhouette scores and stability checks

Text Mining

Corpus statistics without leaving the stack.

  • TF-IDF Vectorization

    n-grams and vocabulary control

  • Topic Modeling

    NMF and LDA over document collections

  • Text Classification

    Supervised labels on corpus features

  • Similarity & Clustering

    Document distance and grouping

Data Engineering

Get data in shape and at scale: SQL over files, streaming frames, and every stats-package format.

SQL Analytics

Full SQL over your files — no database server to stand up.

  • Query Files Directly

    CSV, JSON, and Parquet by path

  • Relational Analytics

    Joins, window functions, and pivots

  • Beyond-RAM Aggregation

    Spills to disk instead of dying

Large-Data Processing

Datasets bigger than memory, handled by design.

  • Streaming DataFrames

    Lazy, out-of-core execution plans

  • Parquet & Arrow

    Columnar interchange between engines

  • Chunked Workflows

    Jobs shaped to the compute envelope

Cleaning & Recoding

Every derivation reproducible, raw data untouched.

  • Coercion & Missing Data

    Tolerant parsing and sentinel handling

  • Reshape & Derive

    Merges, melts, and recodes written as new files

  • Stats-Package Formats

    SPSS, SAS, Stata, plus Excel and ODS

  • Spreadsheet Round-Trip

    Formatted Excel output for collaborators

Visualization & Networks

Publication graphics and graph analysis, saved as first-class workspace artifacts.

Publication Graphics

Every chart a saved script, every figure a durable artifact.

  • Chart Scripts

    PNG, SVG, and PDF exports

  • Statistical Plots

    Box, violin, regression, heatmaps, facets

  • Consistent Styling

    House style kept across a whole report

  • Report-Ready Tables

    Formatted tables for documents and exports

Network & Graph Analysis

The engine behind literature mapping and any relational data.

  • Centrality & Structure

    Degree, betweenness, closeness, PageRank

  • Community Detection

    Modularity-based grouping

  • Science Mapping

    Co-citation, co-authorship, bibliographic coupling

Documents & Media

PDFs, scanned pages, notebooks, and media — read, extract, and produce.

PDF Extraction & OCR

From scanned pages to analysis-ready tables.

  • Text & Table Extraction

    Layout-aware parsing of digital PDFs

  • OCR

    Tesseract on scanned documents and figures

  • Page Rendering

    Page images for figures and inspection

  • Assemble & Split

    Merge, split, and metadata edits

Reproducible Notebooks

The deliverable can be the method.

  • Notebook Authoring

    .ipynb built programmatically

  • Headless Execution

    Execute and re-run notebooks headlessly in isolated analysis jobs

  • Saved Beside Outputs

    Notebook, data, and figures in one workspace

Image & Media Processing

Everyday media work without leaving the workspace.

  • Image Processing

    Resize, convert, and compose

  • Audio & Video

    Conversion, inspection, and frame extraction

The Research Stack

Every package, every version, pinned.

The environment is baked into the sandbox image and read-only at runtime, so a script that ran yesterday runs identically tomorrow. For everyone, on every project.

DataFrames & I/O

pandasv2.3.2Polarsv1.33.1PyArrowv21.0.0DuckDBv1.3.2openpyxlv3.1.5odfpyv1.4.1pyxlsbv1.0.10PyReadStatv1.3.1tabulatev0.9.0

Statistics & ML

NumPyv2.3.3SciPyv1.16.2statsmodelsv0.14.5scikit-learnv1.7.2

Visualization & Networks

Matplotlibv3.10.6Seabornv0.13.2NetworkXv3.5

PDF & Documents

PyMuPDFv1.26.4pdfplumberv0.11.7pypdfv6.0.0python-docxv1.2.0python-pptxv1.0.2XlsxWriterv3.2.9markitdownv0.1.7

Notebooks

nbformatv5.10.4nbclientv0.10.2ipykernelv6.30.1

Images

Pillowv11.3.0

System Tools

FFmpegPopplerTesseract OCRGraphvizripgrepjqgitNode.js 24

Honest boundaries

What we don't do, and what we do instead.

A catalog that promises everything is lying about something. These limits are deliberate, and each has a research-grade path forward.

Deep Learning & GPU Training

The sandbox is CPU-only; PyTorch and TensorFlow are not part of the pinned stack

InsteadClassical ML covers tabular prediction, with small MLP baselines included

XGBoost / LightGBM

Not pinned in the environment

Insteadscikit-learn's HistGradientBoosting — same family, in the box

R, Stata, SPSS, SAS Software

The software isn't installed — only their file formats are readable

InsteadDatasets load natively from .sav, .dta, and .sas7bdat, and the workflow translates to Python

Hours-long Training Jobs

Runs are bounded (~2 min, 2 GB) to keep research turns interactive

InsteadStream with DuckDB or Polars, chunk across runs, or simplify the model

Installing Packages at Runtime

Deliberately disabled so every run is reproducible

InsteadThe stack is curated and version-pinned; request an addition and it ships to every sandbox

Network Access from Analysis Code

Analysis runs network-isolated, without your credentials

InsteadThe assistant fetches data host-side (literature, World Bank, web), then analyzes by path

Bring your dataset.

Upload a CSV, an SPSS file, a folder of PDFs, and ask for the analysis. The assistant picks the method, runs it here, and writes the results back to your workspace.

Open Workspace