The Analysis Capabilities
Real Analysis, on a Real Scientific Stack.
Every capability below runs in an isolated analysis cloud computer with a scientific Python stack and durable project outputs.
16
Analysis Domains
67
Capabilities
35
Research Packages
How Analysis Runs Happen
A lab environment, not a code snippet box.
Each analysis command starts an isolated job, stages the project from private cloud storage, runs without network access, and synchronizes changed outputs back to the project.
Real Cloud Computer
A private microVM per user
Full Python, Node, and shell, not a snippet runner
Research Environment
Read-only, reproducible
Every project runs the same vetted scientific stack; nothing installs at runtime
Compute Envelope
Sufficient RAM
Sized for interactive research turns; heavy jobs chunk or stream
No Network Inside
Your data stays in the sandbox
Bulk collection runs host-side; analysis works on files by path
Provenance by Default
Raw data is immutable
Every derivation writes a new, clearly named file — rerunnable anytime
Saved as Artifacts
Scripts, figures, notebooks
Each run lands in your project workspace next to the data
Statistics & Inference
From first descriptives to models with defensible uncertainty.
Descriptive Statistics
Know your data before modeling it.
Summary Statistics
Means, medians, spread, quartiles — missing values accounted
Frequency Tables
Level counts and percentages, reconciled and sorted
Dataset Profiling
Row/column census with types, missing, distinct, samples
Cross-Tabulation
Contingency tables for categorical relationships
Distribution Shape
Skewness, kurtosis, and quantile checks
Hypothesis Testing
Parametric, non-parametric, and everything the assumptions require.
t-tests & ANOVA
Group comparisons with post-hoc follow-ups
Non-Parametric Tests
Mann-Whitney, Wilcoxon, Kruskal-Wallis, χ², Kolmogorov-Smirnov
Correlation with CIs
Pearson, Spearman, and Kendall with significance
Effect Sizes & Intervals
Estimates with confidence, not just p-values
Power Analysis
Sample sizing and minimum detectable effects
Multiple-Testing Control
Bonferroni, Holm, and FDR corrections
Regression & Generalized Models
Models with full inference — coefficients, intervals, and diagnostics.
Linear Models
OLS/WLS/GLS with R², F-tests, residual diagnostics
Generalized Linear Models
Logistic, Poisson, and negative binomial
Robust Standard Errors
HC estimators and cluster-robust inference
Mixed-Effects Models
Random intercepts and slopes for multilevel data
GEE
Correlated and longitudinal designs
Nonlinear Curve Fitting
Custom models fit with confidence bands
Time Series & Forecasting
Trends, seasonality, and forecasts from your own series.
ARIMA & SARIMAX
Seasonal models with exogenous regressors
Exponential Smoothing
ETS state-space forecasting
Decomposition
Trend, seasonal, and remainder separation
Vector Autoregression
Multi-series dynamics and impulse response
Signal Processing
Filters, spectra, and peak detection
Survival & Duration Analysis
Time-to-event modeling that respects censoring.
Kaplan-Meier
Survival curves with confidence bands
Cox Proportional Hazards
Semiparametric regression with diagnostics
Parametric Survival
Accelerated failure-time families
Machine Learning
Classical, tabular, CPU-scale, the kind of ML research questions actually need.
Predictive Modeling
Train, validate, and persist — the full supervised workflow.
Classification & Regression
Logistic, SVM, k-NN, naive Bayes
Tree Ensembles
Random forests and gradient boosting
Neural Baselines
MLPs for small tabular problems
Cross-Validation
K-fold, stratified, and grouped splits
Hyperparameter Search
Grid and randomized tuning
Evaluation & Calibration
ROC/AUC, PR curves, confusion matrices
Model Persistence
Fitted models saved to your workspace
Unsupervised Learning
Find structure when nobody handed you labels.
Clustering
k-means, DBSCAN, and agglomerative
Gaussian Mixtures
EM-fitted soft clustering
Dimensionality Reduction
PCA, t-SNE, and MDS
Anomaly Detection
Isolation forests and density methods
Cluster Diagnostics
Silhouette scores and stability checks
Text Mining
Corpus statistics without leaving the stack.
TF-IDF Vectorization
n-grams and vocabulary control
Topic Modeling
NMF and LDA over document collections
Text Classification
Supervised labels on corpus features
Similarity & Clustering
Document distance and grouping
Data Engineering
Get data in shape and at scale: SQL over files, streaming frames, and every stats-package format.
SQL Analytics
Full SQL over your files — no database server to stand up.
Query Files Directly
CSV, JSON, and Parquet by path
Relational Analytics
Joins, window functions, and pivots
Beyond-RAM Aggregation
Spills to disk instead of dying
Large-Data Processing
Datasets bigger than memory, handled by design.
Streaming DataFrames
Lazy, out-of-core execution plans
Parquet & Arrow
Columnar interchange between engines
Chunked Workflows
Jobs shaped to the compute envelope
Cleaning & Recoding
Every derivation reproducible, raw data untouched.
Coercion & Missing Data
Tolerant parsing and sentinel handling
Reshape & Derive
Merges, melts, and recodes written as new files
Stats-Package Formats
SPSS, SAS, Stata, plus Excel and ODS
Spreadsheet Round-Trip
Formatted Excel output for collaborators
Visualization & Networks
Publication graphics and graph analysis, saved as first-class workspace artifacts.
Publication Graphics
Every chart a saved script, every figure a durable artifact.
Chart Scripts
PNG, SVG, and PDF exports
Statistical Plots
Box, violin, regression, heatmaps, facets
Consistent Styling
House style kept across a whole report
Report-Ready Tables
Formatted tables for documents and exports
Network & Graph Analysis
The engine behind literature mapping and any relational data.
Centrality & Structure
Degree, betweenness, closeness, PageRank
Community Detection
Modularity-based grouping
Science Mapping
Co-citation, co-authorship, bibliographic coupling
Documents & Media
PDFs, scanned pages, notebooks, and media — read, extract, and produce.
PDF Extraction & OCR
From scanned pages to analysis-ready tables.
Text & Table Extraction
Layout-aware parsing of digital PDFs
OCR
Tesseract on scanned documents and figures
Page Rendering
Page images for figures and inspection
Assemble & Split
Merge, split, and metadata edits
Reproducible Notebooks
The deliverable can be the method.
Notebook Authoring
.ipynb built programmatically
Headless Execution
Execute and re-run notebooks headlessly in isolated analysis jobs
Saved Beside Outputs
Notebook, data, and figures in one workspace
Image & Media Processing
Everyday media work without leaving the workspace.
Image Processing
Resize, convert, and compose
Audio & Video
Conversion, inspection, and frame extraction
The Research Stack
Every package, every version, pinned.
The environment is baked into the sandbox image and read-only at runtime, so a script that ran yesterday runs identically tomorrow. For everyone, on every project.
DataFrames & I/O
Statistics & ML
Visualization & Networks
PDF & Documents
Notebooks
Images
System Tools
Honest boundaries
What we don't do, and what we do instead.
A catalog that promises everything is lying about something. These limits are deliberate, and each has a research-grade path forward.
Deep Learning & GPU Training
The sandbox is CPU-only; PyTorch and TensorFlow are not part of the pinned stack
InsteadClassical ML covers tabular prediction, with small MLP baselines included
XGBoost / LightGBM
Not pinned in the environment
Insteadscikit-learn's HistGradientBoosting — same family, in the box
R, Stata, SPSS, SAS Software
The software isn't installed — only their file formats are readable
InsteadDatasets load natively from .sav, .dta, and .sas7bdat, and the workflow translates to Python
Hours-long Training Jobs
Runs are bounded (~2 min, 2 GB) to keep research turns interactive
InsteadStream with DuckDB or Polars, chunk across runs, or simplify the model
Installing Packages at Runtime
Deliberately disabled so every run is reproducible
InsteadThe stack is curated and version-pinned; request an addition and it ships to every sandbox
Network Access from Analysis Code
Analysis runs network-isolated, without your credentials
InsteadThe assistant fetches data host-side (literature, World Bank, web), then analyzes by path
Bring your dataset.
Upload a CSV, an SPSS file, a folder of PDFs, and ask for the analysis. The assistant picks the method, runs it here, and writes the results back to your workspace.