An elastically-scaled, metric-driven cloud execution platform combining JVM bytecode instrumentation, ML-based request-cost prediction, and a custom Load Balancer / Auto-Scaler to schedule CPU-intensive workloads across AWS EC2 and Lambda
- Overview
- Key Features
- System Architecture
- Instrumentation & Metric Pipeline
- Request Cost Estimation
- Load Balancing & Auto-Scaling
- Project Structure
- Getting Started
- Infrastructure Deployment
- Workload API
- Technologies
- Resources
- Authors
Nature@Cloud is a fully automated elastic cloud service that executes three CPU-bound nature-inspired workloads - Julia-set fractals, Gray-Scott reaction-diffusion, and DNA sequence matching - on a dynamically managed cluster of AWS EC2 workers with serverless AWS Lambda overflow.
What makes the system unusual is the depth of its routing intelligence. Rather than relying solely on reactive CPU-utilisation signals from CloudWatch, the system instruments every request at the JVM bytecode level, derives environment-independent complexity metrics, trains per-workload ML models to predict request cost before execution, converts that prediction into a vCPU-seconds budget estimate, and uses the resulting budget to make proactive routing and scaling decisions - all within a 50 ms latency budget on the critical path.
The three-stage metric-selection pipeline (10 candidate metrics -> 5 filter-stage survivors -> 2 final metrics), the ML predictor stack (Ridge + Histogram Gradient Boosting hybrid, online retraining, rule-based cold-start fallback), and the vCPU-seconds debt tracker are described in detail below.
- Custom Javassist agent (
RoutingMetrics.java) injected at JVM class-load time - zero source-code changes to workloads - Weighted Opcode Count (WOC): JVM analogue of EVM gas - each opcode weighted by approximate CPU-cycle cost (tiers 1/2/3/5/8), derived from Ewasm/EVM cost tables calibrated to Intel cycle counts
- External Work (
external_work): argument-aware cost of uninstrumented JDK calls - discovered necessary when DNAminLengthchanges doubled elapsed time while WOC stayed flat - Single
ThreadLocal<long[2]>, single bytecode pass, ~+111% geometric-mean overhead (down from ~+140% for the two-separate-tool baseline) - Three-stage empirical validation: overhead measurement, pairwise Spearman orthogonality, leave-one-out (LOO) R² drop
- Per-workload hybrid predictor: Ridge regression (global trend + safe extrapolation) + HGB residual model (non-linear interior correction)
- Routing score formula:
score = 0.5153·√woc + 0.4847·√external_work(in-sample R² ≈ 0.987 on n=468 trials) - Predictors served by a Python Flask sidecar on the gateway instance (
127.0.0.1:8081/predict), 50 ms timeout - Rule-based cold-start fallback
score = α·C^βwithβ≈0.5(derived analytically from the sqrt-score formula); per-workload complexity terms fitted on real data
- LB counts completions per workload; at threshold N (default 50, configurable) reads MSS rows, appends (de-duplicated) to dataset CSV, POSTs
/trainto the sidecar - Sidecar retrains on a background thread, atomically hot-swaps in-memory model on success; old model continues serving during training
- Dataset persisted to
analysis/lb_predictors/dataset/; artefacts toanalysis/lb_predictors/ec2/models/
- Routing score -> vCPU-seconds via power-law calibration fitted on isolated t3.micro thread-CPU-time measurements:
cpu = K·score^α(fractals R²=0.993, DNA R²=0.880; GrayScott R²=0.708 - irreducible due to PDE phase transitions) - Per-instance additive debt tracker:
route iff (debt + new_cost) ≤ B_safe = 0.90 × 2 vCPUs × 60 s = 108 vCPU-s - CloudWatch used only as drift correction (EMA-smoothed blend when
|internal − observed| > 0.25·B_safefor two consecutive polls); routing decisions read only internal state
- Event-driven daemon threads; LB and AS communicate via a
BlockingQueue<OrchestrationEvent> - Debt-based least-loaded instance selection; Lambda overflow when no EC2 instance can accept the request
- AS convergence loop: desired =
ceil(total_load / (HIGH_THRESHOLD × B_safe)), clamped to[MIN, MAX]machines - Health monitor with configurable failure/recovery thresholds; draining state prevents mid-request termination
- Silent client-transparent retry on worker failure
GatewayProxyHandlerwraps every request in a resilience loop: on network failure it marks the nodeUNHEALTHY, flushes the routing future, and resubmits to the LB - client sees no error- TTL eviction in
CpuBudgetTrackerprevents ghost debt from crashed workers - Idle hard-flush: debt reset to 0 when CloudWatch reports <2% CPU over two consecutive polls with no in-flight requests
┌──────────────────────────────────────────────────────────────────────┐
│ Client / Browser │
└─────────────────────────────────┬────────────────────────────────────┘
│ HTTP
┌─────────────────────────────────▼────────────────────────────────────┐
│ Gateway EC2 Instance │
│ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ GatewayProxyHandler (retry loop, resilience) │ │
│ └──────────────────────────┬──────────────────────────────────┘ │
│ │ │
│ ┌──────────────────────────▼──────────────────────────────────┐ │
│ │ LoadBalancer (debt-based routing, Lambda overflow) │ │
│ │ ├── ComplexityPredictor -> sidecar :8081/predict │ │
│ │ │ └── RuleBasedPredictor (cold-start fallback) │ │
│ │ ├── CpuBudgetTracker (vCPU-seconds debt per instance) │ │
│ │ └── RetrainManager (threshold -> /train) │ │
│ └───────────┬────────────────────────────┬────────────────────┘ │
│ │ EC2 │ Lambda │
│ ┌───────────▼──────────┐ ┌───────────▼──────────┐ │
│ │ AutoScaler │ │ AWS Lambda │ │
│ │ HealthMonitor │ │ (overflow) │ │
│ │ CpuMonitorThread │ └───────────────────────┘ │
│ │ InstanceManager │ │
│ │ CapacityManager │ │
│ └──────────────────────┘ │
│ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ Python Sidecar (Flask, :8081) │ │
│ │ ├── /predict -> FractalsPredictor / GrayScottPredictor │ │
│ │ │ -> DnaPredictor (two-branch: stop/no-stop) │ │
│ │ └── /train -> async retrain + atomic hot-swap │ │
│ └─────────────────────────────────────────────────────────────┘ │
│ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ DynamoDB (MSS) – metrics + params per completed request │ │
│ └─────────────────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────────────────┘
│ │
┌───────────▼───┐ ┌─────▼───────────┐
│ Worker EC2 │ │ Worker EC2 … │
│ (port 8000) │ │ (port 8000) │
│ + RoutingMetrics Javassist agent │
└───────────────┘ └─────────────────┘
Wall-clock elapsed time is environment-dependent (OS jitter, co-scheduled workloads, I/O). We needed an environment-independent quantity of work that the gateway can use to estimate cost before seeing any execution.
| Metric | Tool | What it counts |
|---|---|---|
insts, blocks, methods |
ICount (teaching team) | bytecodes, basic blocks, method invocations |
branches |
BranchCount | branch instructions |
allocs, alloc_bytes |
AllocationsCount, AllocationSize | NEW/*NEWARRAY count and total bytes |
mac |
MemoryAccessesCount | *ALOAD/*ASTORE, GETFIELD/PUTFIELD |
external_calls |
FanOutCount | calls into uninstrumented (JDK) classes |
external_work |
ExternalCallCount | argument-aware cost of uninstrumented calls |
woc |
WeightedOpcodeCount | tier-weighted opcode sum (EVM gas analogue) |
ICount's three ThreadLocal fields (insts / blocks / methods) caused the highest overhead. We collapsed them into ICountOptimized - a single ThreadLocal<long[3]> - reducing overhead 20–35%. The same single-array pattern was applied to every subsequent tool.
A WOC-fidelity check (WOC vs. ICount in log-log space) showed every workload forming a tight near-linear band, meaning WOC captures the same information as ICount while weighting opcodes by cost. ICount, blocks, and methods were dropped.
The full pairwise Spearman orthogonality matrix (log1p-scaled values, |ρ|>0.90 = redundant) revealed: {woc, insts, blocks, branches} one group; {methods, internal_calls} another; {mac, alloc_bytes} a third. Survivors: woc, mac, allocs, alloc_bytes, external_calls.
When profiling DNA at fixed sequence sizes, doubling minLength doubled elapsed time but woc and external_calls stayed flat. The cost was hidden inside String.equals/substring - uninstrumented JDK methods. Counting calls as one each misses argument magnitude entirely.
We built ExternalCallCount using ExprEditor.edit(MethodCall) to rewrite every external call site: before each call proceeds, it evaluates a static cost expression that depends on the runtime argument sizes (e.g. Math.min(seq1.length(), seq2.length()) for String.equals). The result accumulates as external_work, which now scales linearly with minLength.
Filter decision used two measures:
- Standalone Spearman ρ: does the metric alone correlate with elapsed time?
- Leave-One-Out (LOO) R² drop: how much does the joint ridge regression lose when this metric is removed? Large drop = unique, irreplaceable information.
Results: allocs, alloc_bytes, mac all had LOO < 0.005. external_work subsumed external_calls (near-identical ρ≈0.76, higher LOO). At checkpoint we kept mac as a defensive hedge. Post-checkpoint, teaching team feedback prioritised minimising overhead; with LOO < 0.005, mac was dropped.
A synthetic purity check (four synthetic workloads each designed to isolate one dimension) confirmed instrumentation measures what its names claim.
For each surviving metric, four transforms predicting log1p(elapsed_s) were compared (all R² in the same target space):
woc: linear=0.779, log=0.513, sqrt=0.831, power=0.513external_work: linear=0.619, log=0.645, sqrt=0.811, power=0.645
sqrt wins for both. This is consistent with the power-law exponent β≈0.5 found in the rule-based fallback (see below).
Ridge regression on z-scored sqrt features: in-sample R²≈0.987, n=468.
| Metric | Weight |
|---|---|
woc |
0.5153 |
external_work |
0.4847 |
Final routing formula:
score = 0.5153·√woc + 0.4847·√external_work
RoutingMetrics.java merges both metrics into one Javassist tool:
- One
ThreadLocal<long[2]>(slot 0 = WOC, slot 1 = external_work) - One bytecode transformation pass per class load
The critical subtlety: Phase 1 (ExprEditor) rewrites external call sites on the original bytecode - each INVOKEVIRTUAL is expanded into a cost-charging preamble plus the original call, shifting subsequent bytecode offsets. Phase 2 then computes basic-block positions from the already-modified bytecode and inserts constant WOC probes per block. Running the phases in the wrong order produces stale block positions and incorrect WOC counts. This two-phase approach is why ExprEditor is used for external_work (runtime argument sizes, expression-level rewriting needed) while insertAt suffices for WOC (static pre-summed constant per block, no runtime inspection needed). Result: bit-identical metric values at ~+111% overhead vs. ~+140% for the two-separate-tool baseline.
The routing score is only knowable after execution (when the Javassist agent has produced woc and external_work). The gateway must estimate it from request parameters alone, before routing.
Three separate predictors, one per workload:
Model architecture:
- Ridge regression: linear model, extrapolates safely beyond training range, prevents wild out-of-distribution predictions
- HGB (Histogram Gradient Boosting): captures non-linear interactions inside training range, but clips at training boundaries - dangerous for unusually expensive requests
- Hybrid: Ridge base + HGB fitted on Ridge residuals. Ridge guarantees sane extremes; HGB sharpens interior accuracy
Training target: log1p(score) (stabilises variance across several orders of magnitude). Model selection on tail-aware metric: MAE on the top-10% most expensive requests. Ridge preferred within 5% margin for extrapolation safety.
GrayScott early-termination fix: stopOnExtinction=true cases terminate early; a dedicated Ridge sub-model (_iter_estimator) predicts actual iteration count from (f, k, size, maxIter, stopOnExtinction), which then feeds the complexity feature log1p(size²·estimated_iters). Lifted tail R² from strongly negative to ≈0.92.
DNA worst-case labelling: stopOnFirst=true termination depends on sequence content, which is unobservable at routing time. Since under-estimation is catastrophic for routing (overloads a worker), the dna_true branch is trained on upper-bound labels from the corresponding stopOnFirst=false runs with matching parameters, ensuring safe over-estimates.
Accuracy (held-out test set):
| Workload | MAPE | R² |
|---|---|---|
| fractals | ~0.6% | ~0.997 |
| grayscott | ~6.8% | ~0.968 |
| dna | ~7.1% | ~0.986 |
Required by the project: the system must route before any trained model exists.
Why β≈0.5 analytically:
score = w_woc·√woc + w_ext·√ext
woc ≈ k·C -> score ≈ (w_woc·√k)·√C = α·C^0.5
The sqrt in the score formula maps directly to β=0.5 in the fallback, confirmed empirically:
| Workload | C formula | α | β | R² |
|---|---|---|---|---|
| fractals | w·h·log(1+iter) |
7.629 | 0.515 | 0.988 |
| grayscott | size²·maxIter (stop=False only) |
9.691 | 0.499 | 1.000 |
| dna | seq1·seq2·minLen (stop=False only) |
0.378 | 0.535 | 0.931 |
Fractals: naive C = w·h·iter gave R²=0.33 - Mandelbrot pixels escape early so average work grows logarithmically, not linearly, with the iteration cap. C = w·h·log(1+iter) restores R²=0.988.
GrayScott: fit only on stopOnExtinction=false rows (where actual_iters = maxIter exactly). For stop=true cases this produces a safe over-estimate, which is the correct direction for routing.
DNA: same strategy - fit on stopOnFirst=false, accept safe over-estimates for early-stopping cases.
Constants stored in heuristic_constants.json (bundled in gateway JAR as classpath resource).
Raw score is dimensionless. The gateway needs to answer: "will adding this request push this instance over 90% CPU?"
CPU% and work (score) are different dimensions - you cannot add percentages from different time windows. The bridging currency is vCPU-seconds:
CPU% = (vCPU-seconds / (window × #vCPUs)) × 100
Calibration: power-law cpu_time = K·score^α fitted via OLS through origin on thread-CPU-time measurements (ThreadMXBean.getCurrentThreadCpuTime()) from an isolated t3.micro. The power-law exponent α≈2 emerges structurally: since score ∝ √woc and cpu_time ∝ woc, the true relationship is cpu_time ∝ score².
| Workload | K | α | R² |
|---|---|---|---|
| fractals | 9.22e-10 | 1.964 | 0.993 |
| grayscott | 3.14e-10 | 2.102 | 0.708 |
| dna | 9.45e-07 | 1.308 | 0.880 |
GrayScott's R²=0.708 is irreducible: the Gray-Scott PDE has phase transitions between pattern regimes (spots, waves, stripes, chaos) controlled by (f,k). Two requests with identical routing scores but different (f,k) can differ in CPU time by orders of magnitude - information the score cannot encode. 0.708 is the best achievable from this feature set.
Routing decision (t3.micro, B = 2 vCPUs × 60 s = 120 vCPU-s, B_safe = 108):
Route to instance I iff: debt(I) + K·score^α ≤ 108 vCPU-s
Pick the eligible instance with the lowest current debt. CloudWatch is used only as a slow drift-correction loop - never as the primary routing signal.
- Single daemon thread (
LoadBalancer-Daemon) processes queued requests sequentially under a write lock - Debt-based least-loaded selection: iterate healthy instances, check
canAccept(), pick lowest-debt eligible instance - Lambda overflow: if no EC2 instance can accept, and the request is below a complexity threshold, route to AWS Lambda
- If no routing is possible: park on
routingLock.wait()until a completion or a new instance frees capacity - Symmetric
onRouted/onCompletedcalls to bothCpuBudgetTracker(primary routing signal) andCapacityManager(AS total-load signal)
- Event loop consuming
BlockingQueue<OrchestrationEvent>:LOAD_CHANGED,DRAINING_REQUEST_FINISHED,INSTANCE_UNHEALTHY - Desired instance count:
clamp(ceil(total_load / (HIGH_THRESHOLD × B_safe)), MIN, MAX) - Total load = CapacityManager aggregate (complexity load + EMA-decayed CloudWatch signal) + LB queue backlog
- Scale-up/scale-down via deferred
ScheduledFuturetimers (stabilisation window before acting) - Priority removal order: UNHEALTHY -> INITIALIZING -> idle HEALTHY -> least-loaded HEALTHY (-> DRAINING if in-flight)
- Periodic HTTP probes to
/healthon every registered instance - N consecutive failures ->
UNHEALTHY+INSTANCE_UNHEALTHYevent (triggers replacement) - M consecutive successes from UNHEALTHY ->
HEALTHY(reintegrated into routing pool, routing lock notified)
GatewayProxyHandlerretry loop: network failure -> mark node UNHEALTHY -> flushworkerFuture-> resubmit to LB -> client transparentCpuBudgetTrackerTTL eviction: background cleaner every 10 s evicts expired in-flight entries (max TTL = min(2×expectedDuration, 5 min))- Idle hard-flush: debt reset to 0 when <2% CPU for 2 consecutive CloudWatch polls with 0 in-flight requests
nature-at-cloud/
├── src/
│ ├── gateway/ # Load Balancer, Auto-Scaler, Proxy, Sidecar client
│ │ ├── core/
│ │ │ ├── connections/ # InstanceConnection, InstanceManager
│ │ │ ├── orchestration/ # AutoScaler, CapacityManager, CpuBudgetTracker,
│ │ │ │ # CpuMonitorThread, HealthMonitor, OrchestrationEvent
│ │ │ ├── routing/ # LoadBalancer, ComplexityPredictor, RuleBasedPredictor,
│ │ │ │ # RetrainManager, Request
│ │ │ └── storage/ # MetricsRegistry, MetricsResult, DatabaseCalls,
│ │ │ # DatabaseSyncThread
│ │ ├── handlers/ # GatewayProxyHandler, MetricsHandler, RootHandler
│ │ ├── utils/ # HttpUtils
│ │ ├── GatewayConfig.java # Runtime flag singleton
│ │ ├── GatewayServer.java # Entry point, wiring
│ │ └── resources/
│ │ ├── cpu_calibration.json
│ │ └── heuristic_constants.json
│ ├── worker/ # Workload execution engine (EC2 + Lambda)
│ │ ├── core/
│ │ │ ├── apps/
│ │ │ │ ├── fractals/ # JuliaFractal.java
│ │ │ │ ├── grayscott/ # GrayScott.java
│ │ │ │ └── dna/ # Dna.java, DnaHtmlRenderer.java
│ │ │ └── metrics/ # MetricsManager.java, WorkerMetricsStore.java
│ │ ├── handlers/ # FractalsHandler, GrayScottHandler, DnaHandler,
│ │ │ # HealthHandler, SyntheticHandler, WorkerMetricsHandler
│ │ ├── WorkerServer.java # EC2 entry point
│ │ └── WorkerLambdaRouter.java # Lambda entry point
│ ├── javassist/ # Bytecode instrumentation tools
│ │ └── tools/
│ │ ├── RoutingMetrics.java # FINAL: WOC + external_work, single pass
│ │ ├── WeightedOpcodeCount.java # WOC standalone
│ │ ├── ExternalCallCount.java # external_work standalone
│ │ ├── ExternalCallCostModel.java # Argument-size cost registry
│ │ ├── ICount.java / ICountOptimized.java
│ │ ├── MemoryAccessesCount.java
│ │ ├── AllocationsCount.java / AllocationSize.java
│ │ ├── BranchCount.java / FanOutCount.java
│ │ └── AbstractJavassistTool.java
│ └── metrics/ # Shared InstrumentationTool interface
│
├── analysis/
│ ├── metrics/ # Three-stage metric selection pipeline
│ │ ├── experiment.py # Data capture runner (full/filter/weights modes)
│ │ ├── analysis_overhead.py # Overhead measurement
│ │ ├── analysis_orthogonality.py # Pairwise Spearman correlation
│ │ ├── analysis_woc_fidelity.py # WOC vs ICount linearity
│ │ ├── analysis_filter_decision.py # Standalone ρ + LOO R² drop
│ │ ├── analysis_purity.py # Synthetic workload purity check
│ │ ├── analysis_functional_form.py # Transform selection (linear/log/sqrt/power)
│ │ ├── analysis_weights.py # Ridge regression -> final weights
│ │ └── output/
│ │ ├── full/ # Stage 1 outputs
│ │ ├── filter/ # Stage 2 outputs
│ │ └── weights/ # Stage 3 outputs + routing_weights.json
│ │
│ └── lb_predictors/
│ ├── ec2/
│ │ ├── models/ # Trained artefacts (.pkl), heuristic_constants.json,
│ │ │ # cpu_calibration.json
│ │ └── server/
│ │ └── predict_server.py # Flask sidecar
│ ├── pipelines/
│ │ ├── train.py # Per-workload training orchestrator
│ │ ├── evaluate.py # End-to-end accuracy evaluation
│ │ ├── fit_heuristic_constants.py # Rule-based fallback fitting
│ │ ├── fit_cpu_calibration.py # score -> vCPU-seconds power-law fit
│ │ └── compare_model_with_fallback.py
│ ├── workloads/
│ │ ├── base.py
│ │ ├── fractals.py # FractalsPredictor
│ │ ├── grayscott.py # GrayScottPredictor (+ iter sub-model)
│ │ └── dna.py # DnaPredictor (two-branch, upper-bound labels)
│ ├── dataset/ # Training CSVs per workload
│ └── utils/
│ └── lb_predictor_utils.py # Ridge/HGB/hybrid pipelines, shared helpers
│
└── infra/
├── Makefile # Unified deployment orchestration
├── terraform/ # VPC, security groups, launch templates, ASG
└── scripts/ # AMI build pipeline (worker + gateway)
For an in-depth explanation of configuration matrices, operational properties, inner code paths, or deployment blueprints, each module contains an individual, dedicated README file that explains its responsibilities better:
- Java Sub-Systems & Flags Matrix: src/README.md
- Gateway Proxy, Routing, & Auto-Scaling Rules: src/gateway/README.md
- Worker Architecture & Execution Pipelines: src/worker/README.md
- Infrastructure Orchestration & Base Profiles: infra/README.md
- AMI Packaging & Automated Virtual Machine Creation: infra/scripts/README.md
- Terraform Target Cloud Infrastructure Blueprints: infra/terraform/README.md
| Dependency | Version |
|---|---|
| Java (OpenJDK) | 11+ |
| Maven | 3.8+ |
| Python | 3.10+ |
| AWS CLI | v2 |
| Terraform | 1.5+ |
mvn clean installcd analysis/lb_predictors/ec2/server
pip install -r requirements.txt # first time only
./run.shThe sidecar starts on 127.0.0.1:8081. On first startup, if no trained .pkl artefacts exist yet, it returns 503 on /predict - the gateway falls back to the rule-based predictor automatically until the first retraining cycle completes.
cd src/gateway
mvn exec:java \
-Duse_database=false \
-Duse_cloud=false \
-Duse_record_prediction=false \
-Dretrain_threshold=50 \
-Dretrain_dataset_dir="analysis/lb_predictors/dataset"Without instrumentation (plain execution):
cd src/worker
mvn exec:java -Dport=8001With Javassist bytecode instrumentation (production mode):
cd src/worker
java \
-cp target/worker-1.0.0-SNAPSHOT-jar-with-dependencies.jar \
-Xbootclasspath/a:../JavassistWrapper/target/JavassistWrapper-1.0.0-SNAPSHOT-jar-with-dependencies.jar \
-javaagent:target/worker-1.0.0-SNAPSHOT-jar-with-dependencies.jar=RoutingMetrics:pt.ulisboa.tecnico.cnv.worker.core.apps.fractals,pt.ulisboa.tecnico.cnv.worker.core.apps.grayscott,pt.ulisboa.tecnico.cnv.worker.core.apps.dna:output \
-Dport=8001 \
pt.ulisboa.tecnico.cnv.worker.WorkerServerThe -javaagent flag loads the RoutingMetrics tool and instruments all classes in the three workload packages at class-load time. Metrics are returned per request in the X-Worker-Metrics HTTP response header as JSON {"woc": ..., "external_work": ...}.
Open additional terminals and repeat step 4 with -Dport=8002, -Dport=8003, etc. The gateway's InstanceManager registers workers as they pass health checks on /health.
| JVM Property | Default | Effect |
|---|---|---|
-Duse_database |
false |
Enable DynamoDB persistence for metrics |
-Duse_cloud |
false |
Enable real AWS EC2 provisioning via launch template |
-Duse_record_prediction |
false |
Log predicted vs. actual scores for accuracy telemetry |
-Dworker_launch_template_id |
null |
AWS Launch Template ID for worker provisioning |
-Dretrain_threshold |
50 |
Completed requests per workload before triggering retraining |
-Dretrain_dataset_dir |
analysis/lb_predictors/dataset |
Dataset CSV directory |
Full cloud deployment is automated through a single Makefile at the project root. The pipeline enforces strict sequencing because later steps depend on artefacts produced by earlier ones.
Create a .env file at the project root:
AWS_ACCOUNT_ID=""
AWS_EC2_SSH_KEYPAR_PATH="~/.ssh/id_aws"
AWS_EC2_SSH_PUB_PATH="~/.ssh/id_aws.pub"
AWS_ACCESS_KEY_ID=""
AWS_SECRET_ACCESS_KEY=""Generate an SSH key pair if needed:
ssh-keygen -t rsa -b 4096 -f ~/.ssh/id_awsmake init # init Terraform backend + providers
make infrastructure # deploy VPC, security groups, IAM roles
make build-worker # build worker AMI + register Launch Template
make build-gateway # build gateway AMI (embeds sidecar + trained artefacts)
make compute # deploy gateway EC2 + worker Auto Scaling Group
Teardown:
make destroy # terminate all compute resources
make clean # remove temp_build/ scratch filesFor AMI cleanup (avoids storage charges):
infra/scripts/cleanup-ami.sh worker
infra/scripts/cleanup-ami.sh gatewayThe gateway listens on port 8000. Replace 127.0.0.1:8000 with the gateway's public IP when running on AWS.
curl -s "http://127.0.0.1:8000/fractals?w=800&h=600&iterations=100" \
| awk -F',' '{print $2}' | tr -d '" \n\r' | base64 -d > julia.png| Parameter | Type | Description |
|---|---|---|
w |
int | Image width in pixels |
h |
int | Image height in pixels |
iterations |
int | Max divergence-test iterations per pixel |
curl -s "http://127.0.0.1:8000/grayscott?size=256&maxIterations=5000&f=0.030&k=0.062&stopOnExtinction=false&seedMode=center" \
| awk -F',' '{print $2}' | tr -d '" \n\r' | base64 -d > grayscott.png| Parameter | Type | Values | Description |
|---|---|---|---|
size |
int | 48–1024 | Grid side length (grid is size×size) |
maxIterations |
int | 100–10000 | Upper bound on simulation steps |
f |
float | 0.022, 0.030, 0.230 | Feed rate |
k |
float | 0.051, 0.062 | Kill rate |
stopOnExtinction |
bool | true/false | Early stop if V concentration collapses |
seedMode |
string | center, ring, stripe | Initial V-seeding pattern |
SEQ1="human_HBB:ATGGTGCATCTGACTCCTGAGGAGAAGTCTGCCGTTACTGCCCTGTGGGGCAAGGTG"
SEQ2="chimpanzee_HBB:ATGGTGCACCTGACTCCTGAGGAGAAGTCTGCCGTTACTGCCCTGTGGGGCAAGGTG"
curl "http://127.0.0.1:8000/dna?seq1=$SEQ1&seq2=$SEQ2&minLength=5&stopOnFirst=false"| Parameter | Type | Description |
|---|---|---|
seq1 |
string | name:FASTA_sequence |
seq2 |
string | name:FASTA_sequence |
minLength |
int | Minimum seed window length for match detection |
stopOnFirst |
bool | Return immediately after first match |
| Category | Technology |
|---|---|
| Language | Java 11, Python 3.10+ |
| Build | Maven 3.8+ (multi-module), Terraform 1.5+ |
| Bytecode instrumentation | Javassist |
| ML | scikit-learn (Ridge, HistGradientBoosting), NumPy, pandas |
| Prediction sidecar | Python Flask |
| AWS services | EC2 (t3.micro workers), Lambda, DynamoDB (MSS), CloudWatch (CPU drift correction), IAM |
| Infrastructure as Code | Terraform + shell scripts + Makefile |
- Full Technical Report: Instrumentation pipeline, metric selection, ML prediction design, LB/AS algorithms, vCPU-seconds calibration, fault tolerance mechanisms
- Ewasm Metering - opcode cost table foundation for WOC weights
- Determining Wasm Gas Costs - Intel cycle cost calibration
- Javassist Tutorial - bytecode manipulation API
Instituto Superior Técnico • Cloud Computing and Virtualization • 2025/2026