One assistant
Spark or a 32–72GB GPUA planner breaks down requests. Lightweight specialists start on demand; heavy work queues or uses authorized external services.
01 / PRIVATE ARCHITECTURE
Echo receives your request and stays with it. A control node routes specialist work to your private compute pool.
Conversation · Presence · Review
Plan tasks · Allocate resources · Track progress
Research / Code / Video / Review
Named provider, task, data and budget.
Only approved context leaves after consent.
External workloads are separate from personal Memory.
The companion comes first.
Echo acknowledges you, keeps the conversation moving, and shows what is happening while deeper work continues.
Presence is real-time; intelligence can take time.
Illustrative interaction · product target
“Can you work through these notes?”
Echo
I’m here. Let me work through that.
Local first. Your choice, always.
Personal memory stays local by default. Cloud assistance shares only the context you explicitly approve.
More capacity. The same boundaries.
When outside help would be useful, Echo explains why, shows what would be shared, and asks first. You can always keep the task local.
Named provider and model, why outside expertise would help, and the proposed task.
Messages, excerpts, files, images, and tool context prepared locally. Nothing is pre-uploaded.
Review retention and training terms, charges or a spending cap, and any unknown completion time. Consent applies only to the disclosed task and payload.
Local means your configured device or private deployment. Remote hosting, backups, telemetry, connected apps, and web requests each need clearly disclosed data boundaries; local inference alone does not make those activities offline.
Policy explanation only. These controls do not connect to an AI service or send any context.
Bind consent to a task, provider, disclosed payload, and cost limit. Any expansion requires new consent. Declining or not responding never authorizes upload.
Prepare the smallest useful payload locally. No background pre-upload, full memory export, or automatic fallback to a second provider.
Continue locally where feasible, queue the work, narrow its scope, or explain that it cannot be completed locally.
Cancel unsent work and stop future transmissions. Explain that already transmitted data cannot be recalled by Echo.
Identify when an answer used external assistance and preserve the task’s sharing record.
Disclose the selected provider’s retention and training terms before consent. Never imply zero retention unless the selected service and configuration support it.
Local means your configured device or private deployment. Remote hosting, backups, telemetry, connected apps, and web requests each need clearly disclosed data boundaries; local inference alone does not make those activities offline.
02 / AGENT SCALING
From planning and execution to review, Echo organizes complex work into results you can inspect.
A planner breaks down requests. Lightweight specialists start on demand; heavy work queues or uses authorized external services.
Planning, execution and review have distinct roles. Two GPUs can separate conversation from heavy video jobs.
Specialist tasks can run in parallel and create subtasks within depth, concurrency and budget limits.
Keep interaction services resident and separate from burst capacity for multiple workflows and batch work.
Evaluate larger model services and heavy workloads with more memory and high-speed multi-GPU interconnects.
Running: 3 demo tasks · Queued: 2 demo tasks
Demo data: the planner splits research, code and video tasks. Explore the next step; no real agents execute.
Agent creation is a software capability. GPU count is not an agent count. Recursive work needs depth, concurrency, budget and timeout limits, plus cancellation propagation. Concurrency requires measurement.
03 / VIDEO AI
Echo coordinates storyboards, generation, review and editing to turn an idea into work you can review.
Text, first/last frames and multimodal references; 4–15 seconds, 24 FPS and native stereo audio.
Open Base defaults to a 768p short edge. The full official 2K workflow includes hosted Context-IR and Regenerate-2K.
Official documentation & examples ↗Text-to-video and image-to-video at 720p / 24 FPS.
Configure each variant separately. Do not assume the native audio features of H3.
Official documentation & examples ↗Multimodal world generation and action conditioning; primary generation settings up to 720p / 30 FPS.
For physical AI. Check Edge limitations separately; this is not a real-time avatar service.
Official documentation & examples ↗Native audio, reference controls and enhanced consistency.
Access through cloud services. More local GPUs do not make Veo locally deployable.
Official documentation & examples ↗A 60-second ad is a multi-shot production goal. Offline video is separate from real-time avatars; character consistency and automated review are workflow goals. Linked examples are from official providers, not footage of an integrated NM service.
04 / AUTONOMOUS COMPUTE ECONOMY
Reserve capacity for yourself. Put permitted idle resources to work, and expand beyond local compute within your budget.
External renting is off by default. Reserve interaction capacity and see personal jobs alongside available headroom.
Stop taking new external jobs first. Drain, preempt or wait according to each contract. A UI switch cannot promise instant recovery.
Zero-trustExternal jobs require separate execution environments, networks, storage, credentials and permissions. Clear tenant data afterward and retain audit records.
Reserved 2 GPUs · Personal use 2 GPUs · External use 0 GPUs · Idle 4 GPUs
2 queued · No recovery needed · External budget: $0
Demo data. Resources stay local and nothing is sent externally. Four GPUs represent eligible headroom only.
Offer authorized GPU capacity by the hour
Deliver inference or video tasks
Deliver completed agent outcomes
Three commercial stages, not three markets launching at once. After reserving two GPUs, the remaining six are not a complete eight-GPU HGX. No simulated live income is shown.
05 / CAPABILITY LADDER
Start with DGX Spark for Echo Core, then explore nine dedicated GPU configurations for generation and multi-GPU production. Size for the workload; memory alone does not determine visual quality.
Showing 10 / 10 reference configurations
Shared CPU + GPU memory
Start with a compact private node for everyday assistance, document search, drafting and task planning. Add a dedicated generation node as your workload grows.
128GB unified memory is shared by the CPU, GPU and operating system. Heavy tasks may queue; model fit and application compatibility require validation.
Official specifications ↗1 GPU × 32GB per GPU
Evaluate local generation with quantization, CPU offload or staged loading.
Heavy jobs share resources with Echo. Full component residency is not guaranteed.
Official specifications ↗1 GPU × 48GB per GPU
More loading headroom than 32GB, designed around an optimized pipeline.
Do not assume every component stays in memory at once. Validate the exact model variant.
Official specifications ↗1 GPU × 72GB per GPU
Keep more components resident and reduce transfers between system and GPU memory.
Both 48GB and 72GB variants have 1,344 GB/s bandwidth. More capacity does not imply proportionally faster generation.
Official specifications ↗1 GPU × 96GB per GPU
Evaluate an optimized H3 service together with video pre- and post-processing.
Coexistence with a large language model depends on measured peak memory.
Official specifications ↗2 GPUs × 96GB per GPU
Reserve GPU 0 for interaction and GPU 1 for video generation.
192GB is aggregate capacity. Split services across cards; it is not one unified memory pool.
Official specifications ↗4 GPUs × 96GB per GPU
One GPU can serve interaction and orchestration while three generate shot candidates.
384GB aggregate. Shot concurrency and generation speed require workload benchmarks.
Official specifications ↗8 GPUs × 96GB per GPU
Separate interaction, generation, reference images and post-production to scale workflow throughput.
768GB aggregate. Budget separately for CPU, RAM, storage, power and networking.
Official specifications ↗1 GPU × 288GB per GPU
Evaluate BF16 and larger working sets with less aggressive quantization and offload.
A server resource tier, not a desktop GPU replacement. No guarantee that every workflow fits.
Official specifications ↗8 GPUs × 288GB per GPU
Use high-speed interconnects for multi-GPU work, shot pools, candidate batches and editing.
2,304GB aggregate. Frameworks must explicitly support sharding and parallel execution.
Official specifications ↗06 / INFRASTRUCTURE
Plan models, interaction capacity, storage, networking, power, cooling and maintenance together. Pricing requires a defined configuration and a current quote.
Validate your primary workflow first.
Plan this stage ↗Reserve resources for interaction and creation.
Plan this stage ↗Scale shots, review and post-production.
Plan this stage ↗Validate large working sets, interconnects and whole-node scheduling.
Plan this stage ↗The five product tiers describe service plans; the hardware configurations are references, not one-to-one equivalents. Explore the original comparison, download it or create a local inquiry brief.
Find your Echo.
Every Echo tier can seek outside expertise with your permission. Higher tiers give more of your work room to happen locally, together, and over time.
Proposed range · final capabilities and availability to be validated.
Echo Core · Guided tasks
Start with Echo at your side: remembering what matters, helping with everyday work, and finding a way forward when a task needs more.
Designed forFor individuals who want private everyday assistance in a compact personal system.
Ways to put it to work
Find a decision in your local notes and show the supporting passage.
Draft a brief from approved local information and flag missing context.
Summarize a document, create a draft, and ask before sharing it.
Bring your notes, documents, and everyday context into one companion.
Turn a request into a short, visible sequence of actions.
Wait locally, simplify the task, or approve outside expertise.
Completes short, bounded workflows with checkpoints; proposes extra resources when a task outgrows the local system.
Light background work, with heavier tasks queued.
Expect to consider outside help for difficult reasoning, large workloads, or tight deadlines.
Continue locally where feasible, queue the work, narrow its scope, or explain that it cannot be completed locally.
~128 GB shared accelerator-addressable memory
DGX Spark-class personal appliance
Single integrated CPU/GPU system with shared memory.
Compact desktop appliance · 1 reference node
NVIDIA DGX Spark system overview ↗Hardware references are illustrative, not certified Echo configurations or performance benchmarks. Installed memory is not fully available to models.
Echo Plus · Routine workflows
Give everyday work more room. Echo Plus brings more reasoning and routine workflows onto your own system, with outside help still available when you choose it.
Designed forFor frequent users and creators who want more of their routine work to stay on their own system.
Ways to put it to work
Combine an approved document collection into a sourced working brief.
Iterate on copy, code, and supporting notes in one session.
Organize new local files and prepare a reviewable digest.
Give everyday tasks more local reasoning capacity.
Delegate repeatable work with clear checkpoints.
Work across documents, code, and ideas with more local headroom.
Carries out routine multi-step work within approved tools and scopes, using checkpoints for exceptions.
Routine background workflows alongside conversation.
Consider outside help for exceptional reasoning, specialist models, or workload bursts; no fixed local completion percentage is promised.
Continue locally where feasible, queue the work, narrow its scope, or explain that it cannot be completed locally.
~192–384 GB aggregate VRAM
2–4 × RTX PRO 6000 Blackwell-class GPUs
Validated multi-GPU workstation or server; plan partitioning and GPU communication around the actual PCIe topology.
Workstation or small private server · 1 reference node
NVIDIA RTX PRO 6000 Blackwell Server Edition ↗Hardware references are illustrative, not certified Echo configurations or performance benchmarks. Installed memory is not fully available to models.
Plus and Pro intentionally overlap at 4 × 96 GB. Assign the tier by a validated workload envelope, including reserved presence capacity, solver workload, concurrency, and supported workflow software. A GPU count alone cannot resolve the overlap.
Echo Pro · Delegated projects
Stay in conversation while Echo works through the bigger picture. Pro adds local capacity for deeper analysis, longer projects, and more work happening together.
Designed forFor developers, researchers, and professionals with sustained, demanding personal workloads.
Ways to put it to work
Trace a problem, propose a patch, and run approved checks while you discuss tradeoffs.
Analyze local papers or reports and build a source-linked synthesis.
Prepare drafts and intermediate outputs across a longer scoped workflow.
Follow a demanding project while continuing the conversation.
Keep plans, intermediate results, and next steps together.
Coordinate multiple approved workflows with visible progress.
Plans and executes longer projects within a user-approved scope, using persistent state, checkpoints, and explicit exceptions.
Longer projects and multiple scheduled workflows.
Offer external specialists or overflow when the selected local models, available capacity, or desired turnaround are insufficient.
Continue locally where feasible, queue the work, narrow its scope, or explain that it cannot be completed locally.
~384–768 GB aggregate VRAM
4–8 × RTX PRO 6000 Blackwell-class GPUs
Multi-GPU server with validated PCIe communication paths and explicit workload placement.
Dedicated private compute server · 1 reference node
NVIDIA RTX PRO 6000 Blackwell Server Edition ↗Hardware references are illustrative, not certified Echo configurations or performance benchmarks. Installed memory is not fully available to models.
Plus and Pro intentionally overlap at 4 × 96 GB. Assign the tier by a validated workload envelope, including reserved presence capacity, solver workload, concurrency, and supported workflow software. A GPU count alone cannot resolve the overlap.
Echo Max · Ongoing operations
Build a private foundation for the work that keeps growing. Echo Max gives demanding projects and ongoing workflows room to run, with one familiar companion at the center.
Designed forFor owners who want a personal compute installation built around demanding, continuous workloads.
Ways to put it to work
Keep several local investigations running with source records and periodic briefs.
Run approved analysis, drafting, and code workflows on separate schedules.
Build and refresh a searchable local knowledge collection across substantial private material.
Maintain ongoing work with checkpoints and clear ownership.
Bring demanding analysis to a dedicated private system.
See progress across multiple workers through Echo.
Coordinates persistent project workers and recurring workflows within explicit permissions, schedules, and resource budgets.
Persistent workers across several demanding projects.
Cloud remains optional for distinctive external capability, unavailable models, or bursts beyond the private system’s capacity.
Continue locally where feasible, queue the work, narrow its scope, or explain that it cannot be completed locally.
~1 TB HBM-class; reference total 1,128 GB
8 × H200 in an HGX/DGX-class system
One eight-GPU H200 SXM server with NVLink/NVSwitch fabric.
Private rack installation · 1 reference node
NVIDIA DGX H100/H200 system introduction ↗Hardware references are illustrative, not certified Echo configurations or performance benchmarks. Installed memory is not fully available to models.
Echo Infinite · Custom ongoing operations
Shape a personal AI environment around your ambitions. Echo Infinite offers the most local headroom in the range for demanding projects, persistent workflows, and a deeply provisioned private deployment.
Designed forFor owners building a personal AI compute installation with extensive local workload capacity and custom operations.
Ways to put it to work
Coordinate multiple demanding investigations over private datasets.
Schedule compatible analysis, coding, and generation workloads on dedicated infrastructure.
Keep local knowledge and project state current across a broad set of approved activities.
Provision the most demanding local workloads in the Echo range.
Plan the deployment around your projects and privacy boundaries.
Keep approved workflows coordinated through a single companion.
Coordinates a broad set of persistent, permissioned workflows with custom resource allocation, oversight, and operating policies.
Broad persistent workloads with custom capacity planning.
Outside help stays available for unique services, unavailable models, and overflow. Even this tier can benefit from external expertise with consent.
Continue locally where feasible, queue the work, narrow its scope, or explain that it cannot be completed locally.
~2+ TB HBM-class; reference total 2,304 GB
8 × B300 in a DGX/HGX-class system
One eight-GPU Blackwell Ultra B300 server with a high-bandwidth scale-up fabric.
Personal AI rack infrastructure · 1 reference node
NVIDIA DGX B300 system introduction ↗Hardware references are illustrative, not certified Echo configurations or performance benchmarks. Installed memory is not fully available to models.
Proposed capabilities; final availability and performance depend on the validated configuration. Reference hardware is illustrative. Cloud assistance is optional and may incur separate charges.
More compute increases capacity to act independently within a defined scope; it never grants broader permissions.
The right kind of room.
Compare the intended local work envelope. Responsive presence and control over sharing are common to every tier.
| Designed for | Core | Plus | Pro | Max | Infinite |
|---|---|---|---|---|---|
| Local capability | Assisted local capability | Expanded local capability | Advanced local capability | Dedicated private compute | Flagship private compute |
| Delegated work | Guided tasks | Routine workflows | Delegated projects | Ongoing operations | Custom ongoing operations |
| Background work | Light background work, with heavier tasks queued. | Routine background workflows alongside conversation. | Longer projects and multiple scheduled workflows. | Persistent workers across several demanding projects. | Broad persistent workloads with custom capacity planning. |
| Responsive presence | Same design goal on every tier | Same design goal on every tier | Same design goal on every tier | Same design goal on every tier | Same design goal on every tier |
| Default privacy | Local first; explicit consent for cloud inference | Local first; explicit consent for cloud inference | Local first; explicit consent for cloud inference | Local first; explicit consent for cloud inference | Local first; explicit consent for cloud inference |
| Outside assistance | Optional, with task-specific consent | Optional, with task-specific consent | Optional, with task-specific consent | Optional, with task-specific consent | Optional, with task-specific consent |
Scroll across to compare all five tiers.
| Reference hardware | Core | Plus | Pro | Max | Infinite |
|---|---|---|---|---|---|
| Reference memory | ~128 GB shared accelerator-addressable memory | ~192–384 GB aggregate VRAM | ~384–768 GB aggregate VRAM | ~1 TB HBM-class; reference total 1,128 GB | ~2+ TB HBM-class; reference total 2,304 GB |
| Reference compute | DGX Spark-class personal appliance | 2–4 × RTX PRO 6000 Blackwell-class GPUs | 4–8 × RTX PRO 6000 Blackwell-class GPUs | 8 × H200 in an HGX/DGX-class system | 8 × B300 in a DGX/HGX-class system |
Plus and Pro intentionally overlap at 4 × 96 GB. Assign the tier by a validated workload envelope, including reserved presence capacity, solver workload, concurrency, and supported workflow software. A GPU count alone cannot resolve the overlap.
Aggregate multi-GPU memory is not automatically a single fast memory pool. Fit and performance depend on sharding, quantization, context/KV cache, activations, runtime overhead, and interconnects. Spark’s 128 GB is shared CPU/GPU memory.
NVIDIA DGX Spark system overview ↗07 / TRUST & FAQ
Clear answers on capabilities, privacy and resource use.
This page supports hardware browsing, workflow illustrations and allocation demos. Echo model services, video generation, capacity markets, automatic purchasing and settlement are planned. Hardware tiers are design guidance, not benchmarks. A saved receipt confirms an inquiry submission.
No. Multi-card memory figures are aggregate. PRO 6000 configurations prioritize service separation. Even B300 HGX with NVLink / NVSwitch needs explicit framework support for sharding and parallelism.
No. Evaluate quantization, offload or staged-loading paths on 32–96GB configurations. Feasibility depends on the variant, precision, input size and software version. There is no universal memory threshold for a complete workflow.
The official complete pipeline includes hosted Context-IR and Regenerate-2K, so it is presented as a hybrid service. Before execution, disclose the materials sent, destination and costs. Users can choose to keep work local.
A 60-second ad is a multi-shot generation, review and editing goal, not a single 60-second H3 output. Offline video and low-latency avatars use different service chains. Character consistency and automated review do not succeed on every attempt.
Personal capacity should be reserved. When demand rises, stop accepting external work first, then drain, cancel preemptible work or wait for the contract to expire. Show release timing and alternatives when immediate recovery is unavailable.
Private memory, retrieval libraries and files are excluded from the proposed tenant data flow. External workloads require separate execution environments, networking, storage and credentials. Isolation must be verified in the backend; a diagram alone is not a security guarantee.
There is no real settlement data yet, so this page makes no live-income or payback claims. Future estimates must disclose price, hours, utilization, electricity and other costs. GPU-hours, tasks and finished outcomes are separate commercial stages.
No. Six available GPUs are not a complete eight-GPU node. Whole-node jobs require safely migrating personal services or providing separate reserved capacity within your rules.
Only your entered name, email, intended work, scale and selected configuration are submitted to this website for infrastructure consultation. No personal AI memory or private files are sent. Do not include passwords, keys or sensitive business information. Local brief copying and downloading remain available.
08 / LET’S BUILD YOURS
Tell us what you want Echo to accomplish, and the scale you have in mind.
From research and coding to a private video team. Define your workflow, then choose the hardware that fits.
A saved receipt confirms submission. NM / NM Technology uses your inquiry to respond to your needs. Do not include passwords or sensitive files.