Skip to content

Latest commit

 

History

177 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Nature@Cloud

Grade: 19.5/20

An elastically-scaled, metric-driven cloud execution platform combining JVM bytecode instrumentation, ML-based request-cost prediction, and a custom Load Balancer / Auto-Scaler to schedule CPU-intensive workloads across AWS EC2 and Lambda


Table of Contents


Overview

Nature@Cloud is a fully automated elastic cloud service that executes three CPU-bound nature-inspired workloads - Julia-set fractals, Gray-Scott reaction-diffusion, and DNA sequence matching - on a dynamically managed cluster of AWS EC2 workers with serverless AWS Lambda overflow.

What makes the system unusual is the depth of its routing intelligence. Rather than relying solely on reactive CPU-utilisation signals from CloudWatch, the system instruments every request at the JVM bytecode level, derives environment-independent complexity metrics, trains per-workload ML models to predict request cost before execution, converts that prediction into a vCPU-seconds budget estimate, and uses the resulting budget to make proactive routing and scaling decisions - all within a 50 ms latency budget on the critical path.

The three-stage metric-selection pipeline (10 candidate metrics -> 5 filter-stage survivors -> 2 final metrics), the ML predictor stack (Ridge + Histogram Gradient Boosting hybrid, online retraining, rule-based cold-start fallback), and the vCPU-seconds debt tracker are described in detail below.


Key Features

Bytecode Instrumentation

  • Custom Javassist agent (RoutingMetrics.java) injected at JVM class-load time - zero source-code changes to workloads
  • Weighted Opcode Count (WOC): JVM analogue of EVM gas - each opcode weighted by approximate CPU-cycle cost (tiers 1/2/3/5/8), derived from Ewasm/EVM cost tables calibrated to Intel cycle counts
  • External Work (external_work): argument-aware cost of uninstrumented JDK calls - discovered necessary when DNA minLength changes doubled elapsed time while WOC stayed flat
  • Single ThreadLocal<long[2]>, single bytecode pass, ~+111% geometric-mean overhead (down from ~+140% for the two-separate-tool baseline)
  • Three-stage empirical validation: overhead measurement, pairwise Spearman orthogonality, leave-one-out (LOO) R² drop

ML-Based Request Cost Prediction

  • Per-workload hybrid predictor: Ridge regression (global trend + safe extrapolation) + HGB residual model (non-linear interior correction)
  • Routing score formula: score = 0.5153·√woc + 0.4847·√external_work (in-sample R² ≈ 0.987 on n=468 trials)
  • Predictors served by a Python Flask sidecar on the gateway instance (127.0.0.1:8081/predict), 50 ms timeout
  • Rule-based cold-start fallback score = α·C^β with β≈0.5 (derived analytically from the sqrt-score formula); per-workload complexity terms fitted on real data

Online Retraining Loop

  • LB counts completions per workload; at threshold N (default 50, configurable) reads MSS rows, appends (de-duplicated) to dataset CSV, POSTs /train to the sidecar
  • Sidecar retrains on a background thread, atomically hot-swaps in-memory model on success; old model continues serving during training
  • Dataset persisted to analysis/lb_predictors/dataset/; artefacts to analysis/lb_predictors/ec2/models/

vCPU-Seconds Capacity Routing

  • Routing score -> vCPU-seconds via power-law calibration fitted on isolated t3.micro thread-CPU-time measurements: cpu = K·score^α (fractals R²=0.993, DNA R²=0.880; GrayScott R²=0.708 - irreducible due to PDE phase transitions)
  • Per-instance additive debt tracker: route iff (debt + new_cost) ≤ B_safe = 0.90 × 2 vCPUs × 60 s = 108 vCPU-s
  • CloudWatch used only as drift correction (EMA-smoothed blend when |internal − observed| > 0.25·B_safe for two consecutive polls); routing decisions read only internal state

Custom Load Balancer & Auto-Scaler

  • Event-driven daemon threads; LB and AS communicate via a BlockingQueue<OrchestrationEvent>
  • Debt-based least-loaded instance selection; Lambda overflow when no EC2 instance can accept the request
  • AS convergence loop: desired = ceil(total_load / (HIGH_THRESHOLD × B_safe)), clamped to [MIN, MAX] machines
  • Health monitor with configurable failure/recovery thresholds; draining state prevents mid-request termination
  • Silent client-transparent retry on worker failure

Fault Tolerance

  • GatewayProxyHandler wraps every request in a resilience loop: on network failure it marks the node UNHEALTHY, flushes the routing future, and resubmits to the LB - client sees no error
  • TTL eviction in CpuBudgetTracker prevents ghost debt from crashed workers
  • Idle hard-flush: debt reset to 0 when CloudWatch reports <2% CPU over two consecutive polls with no in-flight requests

System Architecture

┌──────────────────────────────────────────────────────────────────────┐
│                          Client / Browser                            │
└─────────────────────────────────┬────────────────────────────────────┘
                                  │ HTTP
┌─────────────────────────────────▼────────────────────────────────────┐
│                        Gateway EC2 Instance                          │
│                                                                      │
│  ┌─────────────────────────────────────────────────────────────┐     │
│  │  GatewayProxyHandler  (retry loop, resilience)              │     │
│  └──────────────────────────┬──────────────────────────────────┘     │
│                             │                                        │
│  ┌──────────────────────────▼──────────────────────────────────┐     │
│  │  LoadBalancer  (debt-based routing, Lambda overflow)        │     │
│  │    ├── ComplexityPredictor -> sidecar :8081/predict         │     │
│  │    │     └── RuleBasedPredictor  (cold-start fallback)      │     │
│  │    ├── CpuBudgetTracker  (vCPU-seconds debt per instance)   │     │
│  │    └── RetrainManager  (threshold -> /train)                │     │
│  └───────────┬────────────────────────────┬────────────────────┘     │
│              │ EC2                        │ Lambda                   │
│  ┌───────────▼──────────┐     ┌───────────▼──────────┐               │
│  │   AutoScaler         │     │   AWS Lambda          │              │
│  │   HealthMonitor      │     │   (overflow)          │              │
│  │   CpuMonitorThread   │     └───────────────────────┘              │
│  │   InstanceManager    │                                            │
│  │   CapacityManager    │                                            │
│  └──────────────────────┘                                            │
│                                                                      │
│  ┌─────────────────────────────────────────────────────────────┐     │
│  │  Python Sidecar  (Flask, :8081)                             │     │
│  │    ├── /predict  -> FractalsPredictor / GrayScottPredictor  │     │
│  │    │             -> DnaPredictor (two-branch: stop/no-stop) │     │
│  │    └── /train   -> async retrain + atomic hot-swap          │     │
│  └─────────────────────────────────────────────────────────────┘     │
│                                                                      │
│  ┌─────────────────────────────────────────────────────────────┐     │
│  │  DynamoDB (MSS)  – metrics + params per completed request   │     │
│  └─────────────────────────────────────────────────────────────┘     │
└──────────────────────────────────────────────────────────────────────┘
                         │               │
             ┌───────────▼───┐     ┌─────▼───────────┐
             │  Worker EC2   │     │  Worker EC2  …  │
             │  (port 8000)  │     │  (port 8000)    │
             │  + RoutingMetrics Javassist agent     │
             └───────────────┘     └─────────────────┘

Instrumentation & Metric Pipeline

Why bytecode instrumentation?

Wall-clock elapsed time is environment-dependent (OS jitter, co-scheduled workloads, I/O). We needed an environment-independent quantity of work that the gateway can use to estimate cost before seeing any execution.

Candidate metrics (10 total)

Metric Tool What it counts
insts, blocks, methods ICount (teaching team) bytecodes, basic blocks, method invocations
branches BranchCount branch instructions
allocs, alloc_bytes AllocationsCount, AllocationSize NEW/*NEWARRAY count and total bytes
mac MemoryAccessesCount *ALOAD/*ASTORE, GETFIELD/PUTFIELD
external_calls FanOutCount calls into uninstrumented (JDK) classes
external_work ExternalCallCount argument-aware cost of uninstrumented calls
woc WeightedOpcodeCount tier-weighted opcode sum (EVM gas analogue)

Stage 1 - Full evaluation

ICount's three ThreadLocal fields (insts / blocks / methods) caused the highest overhead. We collapsed them into ICountOptimized - a single ThreadLocal<long[3]> - reducing overhead 20–35%. The same single-array pattern was applied to every subsequent tool.

A WOC-fidelity check (WOC vs. ICount in log-log space) showed every workload forming a tight near-linear band, meaning WOC captures the same information as ICount while weighting opcodes by cost. ICount, blocks, and methods were dropped.

The full pairwise Spearman orthogonality matrix (log1p-scaled values, |ρ|>0.90 = redundant) revealed: {woc, insts, blocks, branches} one group; {methods, internal_calls} another; {mac, alloc_bytes} a third. Survivors: woc, mac, allocs, alloc_bytes, external_calls.

Stage 2 - DNA blind spot & filter decision

When profiling DNA at fixed sequence sizes, doubling minLength doubled elapsed time but woc and external_calls stayed flat. The cost was hidden inside String.equals/substring - uninstrumented JDK methods. Counting calls as one each misses argument magnitude entirely.

We built ExternalCallCount using ExprEditor.edit(MethodCall) to rewrite every external call site: before each call proceeds, it evaluates a static cost expression that depends on the runtime argument sizes (e.g. Math.min(seq1.length(), seq2.length()) for String.equals). The result accumulates as external_work, which now scales linearly with minLength.

Filter decision used two measures:

  • Standalone Spearman ρ: does the metric alone correlate with elapsed time?
  • Leave-One-Out (LOO) R² drop: how much does the joint ridge regression lose when this metric is removed? Large drop = unique, irreplaceable information.

Results: allocs, alloc_bytes, mac all had LOO < 0.005. external_work subsumed external_calls (near-identical ρ≈0.76, higher LOO). At checkpoint we kept mac as a defensive hedge. Post-checkpoint, teaching team feedback prioritised minimising overhead; with LOO < 0.005, mac was dropped.

A synthetic purity check (four synthetic workloads each designed to isolate one dimension) confirmed instrumentation measures what its names claim.

Stage 3 - Functional form and weights

For each surviving metric, four transforms predicting log1p(elapsed_s) were compared (all R² in the same target space):

  • woc: linear=0.779, log=0.513, sqrt=0.831, power=0.513
  • external_work: linear=0.619, log=0.645, sqrt=0.811, power=0.645

sqrt wins for both. This is consistent with the power-law exponent β≈0.5 found in the rule-based fallback (see below).

Ridge regression on z-scored sqrt features: in-sample R²≈0.987, n=468.

Metric Weight
woc 0.5153
external_work 0.4847

Final routing formula:

score = 0.5153·√woc + 0.4847·√external_work

RoutingMetrics - single-pass consolidation

RoutingMetrics.java merges both metrics into one Javassist tool:

  • One ThreadLocal<long[2]> (slot 0 = WOC, slot 1 = external_work)
  • One bytecode transformation pass per class load

The critical subtlety: Phase 1 (ExprEditor) rewrites external call sites on the original bytecode - each INVOKEVIRTUAL is expanded into a cost-charging preamble plus the original call, shifting subsequent bytecode offsets. Phase 2 then computes basic-block positions from the already-modified bytecode and inserts constant WOC probes per block. Running the phases in the wrong order produces stale block positions and incorrect WOC counts. This two-phase approach is why ExprEditor is used for external_work (runtime argument sizes, expression-level rewriting needed) while insertAt suffices for WOC (static pre-summed constant per block, no runtime inspection needed). Result: bit-identical metric values at ~+111% overhead vs. ~+140% for the two-separate-tool baseline.


Request Cost Estimation

Prediction problem

The routing score is only knowable after execution (when the Javassist agent has produced woc and external_work). The gateway must estimate it from request parameters alone, before routing.

Per-workload hybrid ML predictor

Three separate predictors, one per workload:

Model architecture:

  • Ridge regression: linear model, extrapolates safely beyond training range, prevents wild out-of-distribution predictions
  • HGB (Histogram Gradient Boosting): captures non-linear interactions inside training range, but clips at training boundaries - dangerous for unusually expensive requests
  • Hybrid: Ridge base + HGB fitted on Ridge residuals. Ridge guarantees sane extremes; HGB sharpens interior accuracy

Training target: log1p(score) (stabilises variance across several orders of magnitude). Model selection on tail-aware metric: MAE on the top-10% most expensive requests. Ridge preferred within 5% margin for extrapolation safety.

GrayScott early-termination fix: stopOnExtinction=true cases terminate early; a dedicated Ridge sub-model (_iter_estimator) predicts actual iteration count from (f, k, size, maxIter, stopOnExtinction), which then feeds the complexity feature log1p(size²·estimated_iters). Lifted tail R² from strongly negative to ≈0.92.

DNA worst-case labelling: stopOnFirst=true termination depends on sequence content, which is unobservable at routing time. Since under-estimation is catastrophic for routing (overloads a worker), the dna_true branch is trained on upper-bound labels from the corresponding stopOnFirst=false runs with matching parameters, ensuring safe over-estimates.

Accuracy (held-out test set):

Workload MAPE R²
fractals ~0.6% ~0.997
grayscott ~6.8% ~0.968
dna ~7.1% ~0.986

Rule-based cold-start fallback

Required by the project: the system must route before any trained model exists.

Why β≈0.5 analytically:

score = w_woc·√woc + w_ext·√ext
woc ≈ k·C  ->  score ≈ (w_woc·√k)·√C = α·C^0.5

The sqrt in the score formula maps directly to β=0.5 in the fallback, confirmed empirically:

Workload C formula α β R²
fractals w·h·log(1+iter) 7.629 0.515 0.988
grayscott size²·maxIter (stop=False only) 9.691 0.499 1.000
dna seq1·seq2·minLen (stop=False only) 0.378 0.535 0.931

Fractals: naive C = w·h·iter gave R²=0.33 - Mandelbrot pixels escape early so average work grows logarithmically, not linearly, with the iteration cap. C = w·h·log(1+iter) restores R²=0.988.

GrayScott: fit only on stopOnExtinction=false rows (where actual_iters = maxIter exactly). For stop=true cases this produces a safe over-estimate, which is the correct direction for routing.

DNA: same strategy - fit on stopOnFirst=false, accept safe over-estimates for early-stopping cases.

Constants stored in heuristic_constants.json (bundled in gateway JAR as classpath resource).

vCPU-seconds capacity accounting

Raw score is dimensionless. The gateway needs to answer: "will adding this request push this instance over 90% CPU?"

CPU% and work (score) are different dimensions - you cannot add percentages from different time windows. The bridging currency is vCPU-seconds:

CPU% = (vCPU-seconds / (window × #vCPUs)) × 100

Calibration: power-law cpu_time = K·score^α fitted via OLS through origin on thread-CPU-time measurements (ThreadMXBean.getCurrentThreadCpuTime()) from an isolated t3.micro. The power-law exponent α≈2 emerges structurally: since score ∝ √woc and cpu_time ∝ woc, the true relationship is cpu_time ∝ score².

Workload K α R²
fractals 9.22e-10 1.964 0.993
grayscott 3.14e-10 2.102 0.708
dna 9.45e-07 1.308 0.880

GrayScott's R²=0.708 is irreducible: the Gray-Scott PDE has phase transitions between pattern regimes (spots, waves, stripes, chaos) controlled by (f,k). Two requests with identical routing scores but different (f,k) can differ in CPU time by orders of magnitude - information the score cannot encode. 0.708 is the best achievable from this feature set.

Routing decision (t3.micro, B = 2 vCPUs × 60 s = 120 vCPU-s, B_safe = 108):

Route to instance I iff: debt(I) + K·score^α ≤ 108 vCPU-s

Pick the eligible instance with the lowest current debt. CloudWatch is used only as a slow drift-correction loop - never as the primary routing signal.


Load Balancing & Auto-Scaling

Load Balancer

  • Single daemon thread (LoadBalancer-Daemon) processes queued requests sequentially under a write lock
  • Debt-based least-loaded selection: iterate healthy instances, check canAccept(), pick lowest-debt eligible instance
  • Lambda overflow: if no EC2 instance can accept, and the request is below a complexity threshold, route to AWS Lambda
  • If no routing is possible: park on routingLock.wait() until a completion or a new instance frees capacity
  • Symmetric onRouted/onCompleted calls to both CpuBudgetTracker (primary routing signal) and CapacityManager (AS total-load signal)

Auto-Scaler

  • Event loop consuming BlockingQueue<OrchestrationEvent>: LOAD_CHANGED, DRAINING_REQUEST_FINISHED, INSTANCE_UNHEALTHY
  • Desired instance count: clamp(ceil(total_load / (HIGH_THRESHOLD × B_safe)), MIN, MAX)
  • Total load = CapacityManager aggregate (complexity load + EMA-decayed CloudWatch signal) + LB queue backlog
  • Scale-up/scale-down via deferred ScheduledFuture timers (stabilisation window before acting)
  • Priority removal order: UNHEALTHY -> INITIALIZING -> idle HEALTHY -> least-loaded HEALTHY (-> DRAINING if in-flight)

Health Monitor

  • Periodic HTTP probes to /health on every registered instance
  • N consecutive failures -> UNHEALTHY + INSTANCE_UNHEALTHY event (triggers replacement)
  • M consecutive successes from UNHEALTHY -> HEALTHY (reintegrated into routing pool, routing lock notified)

Fault tolerance

  • GatewayProxyHandler retry loop: network failure -> mark node UNHEALTHY -> flush workerFuture -> resubmit to LB -> client transparent
  • CpuBudgetTracker TTL eviction: background cleaner every 10 s evicts expired in-flight entries (max TTL = min(2×expectedDuration, 5 min))
  • Idle hard-flush: debt reset to 0 when <2% CPU for 2 consecutive CloudWatch polls with 0 in-flight requests

Project Structure

nature-at-cloud/
├── src/
│   ├── gateway/          # Load Balancer, Auto-Scaler, Proxy, Sidecar client
│   │   ├── core/
│   │   │   ├── connections/      # InstanceConnection, InstanceManager
│   │   │   ├── orchestration/    # AutoScaler, CapacityManager, CpuBudgetTracker,
│   │   │   │                     # CpuMonitorThread, HealthMonitor, OrchestrationEvent
│   │   │   ├── routing/          # LoadBalancer, ComplexityPredictor, RuleBasedPredictor,
│   │   │   │                     # RetrainManager, Request
│   │   │   └── storage/          # MetricsRegistry, MetricsResult, DatabaseCalls,
│   │   │                         # DatabaseSyncThread
│   │   ├── handlers/             # GatewayProxyHandler, MetricsHandler, RootHandler
│   │   ├── utils/                # HttpUtils
│   │   ├── GatewayConfig.java    # Runtime flag singleton
│   │   ├── GatewayServer.java    # Entry point, wiring
│   │   └── resources/
│   │       ├── cpu_calibration.json
│   │       └── heuristic_constants.json
│   ├── worker/           # Workload execution engine (EC2 + Lambda)
│   │   ├── core/
│   │   │   ├── apps/
│   │   │   │   ├── fractals/     # JuliaFractal.java
│   │   │   │   ├── grayscott/    # GrayScott.java
│   │   │   │   └── dna/          # Dna.java, DnaHtmlRenderer.java
│   │   │   └── metrics/          # MetricsManager.java, WorkerMetricsStore.java
│   │   ├── handlers/             # FractalsHandler, GrayScottHandler, DnaHandler,
│   │   │                         # HealthHandler, SyntheticHandler, WorkerMetricsHandler
│   │   ├── WorkerServer.java     # EC2 entry point
│   │   └── WorkerLambdaRouter.java  # Lambda entry point
│   ├── javassist/        # Bytecode instrumentation tools
│   │   └── tools/
│   │       ├── RoutingMetrics.java        # FINAL: WOC + external_work, single pass
│   │       ├── WeightedOpcodeCount.java   # WOC standalone
│   │       ├── ExternalCallCount.java     # external_work standalone
│   │       ├── ExternalCallCostModel.java # Argument-size cost registry
│   │       ├── ICount.java / ICountOptimized.java
│   │       ├── MemoryAccessesCount.java
│   │       ├── AllocationsCount.java / AllocationSize.java
│   │       ├── BranchCount.java / FanOutCount.java
│   │       └── AbstractJavassistTool.java
│   └── metrics/          # Shared InstrumentationTool interface
│
├── analysis/
│   ├── metrics/          # Three-stage metric selection pipeline
│   │   ├── experiment.py               # Data capture runner (full/filter/weights modes)
│   │   ├── analysis_overhead.py        # Overhead measurement
│   │   ├── analysis_orthogonality.py   # Pairwise Spearman correlation
│   │   ├── analysis_woc_fidelity.py    # WOC vs ICount linearity
│   │   ├── analysis_filter_decision.py # Standalone ρ + LOO R² drop
│   │   ├── analysis_purity.py          # Synthetic workload purity check
│   │   ├── analysis_functional_form.py # Transform selection (linear/log/sqrt/power)
│   │   ├── analysis_weights.py         # Ridge regression -> final weights
│   │   └── output/
│   │       ├── full/     # Stage 1 outputs
│   │       ├── filter/   # Stage 2 outputs
│   │       └── weights/  # Stage 3 outputs + routing_weights.json
│   │
│   └── lb_predictors/
│       ├── ec2/
│       │   ├── models/               # Trained artefacts (.pkl), heuristic_constants.json,
│       │   │                         # cpu_calibration.json
│       │   └── server/
│       │       └── predict_server.py # Flask sidecar
│       ├── pipelines/
│       │   ├── train.py              # Per-workload training orchestrator
│       │   ├── evaluate.py           # End-to-end accuracy evaluation
│       │   ├── fit_heuristic_constants.py  # Rule-based fallback fitting
│       │   ├── fit_cpu_calibration.py      # score -> vCPU-seconds power-law fit
│       │   └── compare_model_with_fallback.py
│       ├── workloads/
│       │   ├── base.py
│       │   ├── fractals.py           # FractalsPredictor
│       │   ├── grayscott.py          # GrayScottPredictor (+ iter sub-model)
│       │   └── dna.py                # DnaPredictor (two-branch, upper-bound labels)
│       ├── dataset/                  # Training CSVs per workload
│       └── utils/
│           └── lb_predictor_utils.py # Ridge/HGB/hybrid pipelines, shared helpers
│
└── infra/
    ├── Makefile                      # Unified deployment orchestration
    ├── terraform/                    # VPC, security groups, launch templates, ASG
    └── scripts/                      # AMI build pipeline (worker + gateway)

Getting Started

Important Note - Dedicated Sub-Module Guides

For an in-depth explanation of configuration matrices, operational properties, inner code paths, or deployment blueprints, each module contains an individual, dedicated README file that explains its responsibilities better:

Prerequisites

Dependency Version
Java (OpenJDK) 11+
Maven 3.8+
Python 3.10+
AWS CLI v2
Terraform 1.5+

1. Build all modules

mvn clean install

2. Start the prediction sidecar

cd analysis/lb_predictors/ec2/server
pip install -r requirements.txt   # first time only
./run.sh

The sidecar starts on 127.0.0.1:8081. On first startup, if no trained .pkl artefacts exist yet, it returns 503 on /predict - the gateway falls back to the rule-based predictor automatically until the first retraining cycle completes.

3. Start the gateway (local mode)

cd src/gateway
mvn exec:java \
  -Duse_database=false \
  -Duse_cloud=false \
  -Duse_record_prediction=false \
  -Dretrain_threshold=50 \
  -Dretrain_dataset_dir="analysis/lb_predictors/dataset"

4. Start a worker

Without instrumentation (plain execution):

cd src/worker
mvn exec:java -Dport=8001

With Javassist bytecode instrumentation (production mode):

cd src/worker
java \
  -cp target/worker-1.0.0-SNAPSHOT-jar-with-dependencies.jar \
  -Xbootclasspath/a:../JavassistWrapper/target/JavassistWrapper-1.0.0-SNAPSHOT-jar-with-dependencies.jar \
  -javaagent:target/worker-1.0.0-SNAPSHOT-jar-with-dependencies.jar=RoutingMetrics:pt.ulisboa.tecnico.cnv.worker.core.apps.fractals,pt.ulisboa.tecnico.cnv.worker.core.apps.grayscott,pt.ulisboa.tecnico.cnv.worker.core.apps.dna:output \
  -Dport=8001 \
  pt.ulisboa.tecnico.cnv.worker.WorkerServer

The -javaagent flag loads the RoutingMetrics tool and instruments all classes in the three workload packages at class-load time. Metrics are returned per request in the X-Worker-Metrics HTTP response header as JSON {"woc": ..., "external_work": ...}.

5. Multi-worker local cluster

Open additional terminals and repeat step 4 with -Dport=8002, -Dport=8003, etc. The gateway's InstanceManager registers workers as they pass health checks on /health.

Gateway runtime flags

JVM Property Default Effect
-Duse_database false Enable DynamoDB persistence for metrics
-Duse_cloud false Enable real AWS EC2 provisioning via launch template
-Duse_record_prediction false Log predicted vs. actual scores for accuracy telemetry
-Dworker_launch_template_id null AWS Launch Template ID for worker provisioning
-Dretrain_threshold 50 Completed requests per workload before triggering retraining
-Dretrain_dataset_dir analysis/lb_predictors/dataset Dataset CSV directory

Infrastructure Deployment

Full cloud deployment is automated through a single Makefile at the project root. The pipeline enforces strict sequencing because later steps depend on artefacts produced by earlier ones.

Prerequisites

Create a .env file at the project root:

AWS_ACCOUNT_ID=""
AWS_EC2_SSH_KEYPAR_PATH="~/.ssh/id_aws"
AWS_EC2_SSH_PUB_PATH="~/.ssh/id_aws.pub"
AWS_ACCESS_KEY_ID=""
AWS_SECRET_ACCESS_KEY=""

Generate an SSH key pair if needed:

ssh-keygen -t rsa -b 4096 -f ~/.ssh/id_aws

Deployment pipeline

make init           # init Terraform backend + providers
make infrastructure # deploy VPC, security groups, IAM roles
make build-worker   # build worker AMI + register Launch Template
make build-gateway  # build gateway AMI (embeds sidecar + trained artefacts)
make compute        # deploy gateway EC2 + worker Auto Scaling Group

Teardown:

make destroy        # terminate all compute resources
make clean          # remove temp_build/ scratch files

For AMI cleanup (avoids storage charges):

infra/scripts/cleanup-ami.sh worker
infra/scripts/cleanup-ami.sh gateway

Workload API

The gateway listens on port 8000. Replace 127.0.0.1:8000 with the gateway's public IP when running on AWS.

Fractals

curl -s "http://127.0.0.1:8000/fractals?w=800&h=600&iterations=100" \
  | awk -F',' '{print $2}' | tr -d '" \n\r' | base64 -d > julia.png
Parameter Type Description
w int Image width in pixels
h int Image height in pixels
iterations int Max divergence-test iterations per pixel

GrayScott

curl -s "http://127.0.0.1:8000/grayscott?size=256&maxIterations=5000&f=0.030&k=0.062&stopOnExtinction=false&seedMode=center" \
  | awk -F',' '{print $2}' | tr -d '" \n\r' | base64 -d > grayscott.png
Parameter Type Values Description
size int 48–1024 Grid side length (grid is size×size)
maxIterations int 100–10000 Upper bound on simulation steps
f float 0.022, 0.030, 0.230 Feed rate
k float 0.051, 0.062 Kill rate
stopOnExtinction bool true/false Early stop if V concentration collapses
seedMode string center, ring, stripe Initial V-seeding pattern

DNA Sequence Matching

SEQ1="human_HBB:ATGGTGCATCTGACTCCTGAGGAGAAGTCTGCCGTTACTGCCCTGTGGGGCAAGGTG"
SEQ2="chimpanzee_HBB:ATGGTGCACCTGACTCCTGAGGAGAAGTCTGCCGTTACTGCCCTGTGGGGCAAGGTG"
curl "http://127.0.0.1:8000/dna?seq1=$SEQ1&seq2=$SEQ2&minLength=5&stopOnFirst=false"
Parameter Type Description
seq1 string name:FASTA_sequence
seq2 string name:FASTA_sequence
minLength int Minimum seed window length for match detection
stopOnFirst bool Return immediately after first match

Technologies

Category Technology
Language Java 11, Python 3.10+
Build Maven 3.8+ (multi-module), Terraform 1.5+
Bytecode instrumentation Javassist
ML scikit-learn (Ridge, HistGradientBoosting), NumPy, pandas
Prediction sidecar Python Flask
AWS services EC2 (t3.micro workers), Lambda, DynamoDB (MSS), CloudWatch (CPU drift correction), IAM
Infrastructure as Code Terraform + shell scripts + Makefile

Resources

Project Report

  • Full Technical Report: Instrumentation pipeline, metric selection, ML prediction design, LB/AS algorithms, vCPU-seconds calibration, fault tolerance mechanisms

Key References


Authors

Tomás Fernandes
Tomás Fernandes
Martim Sousa
Martim Sousa
Cristiano Pantea
Cristiano Pantea

Instituto Superior Técnico • Cloud Computing and Virtualization • 2025/2026

About

An elastically-scaled, metric-driven cloud execution platform combining JVM bytecode instrumentation, ML-based request-cost prediction, and a custom Load Balancer / Auto-Scaler to schedule CPU-intensive workloads across AWS EC2 and Lambda

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages