An elegant silver-haired AI woman with a warm gaze and soft emerald light

One Echo.
Your own AI team.

NM Personal Data Center gives Echo private compute to stay with you,
organize research, coding and creative work, and allocate resources by your rules.

Move your pointer over Echo to dissolve her into light, or choose Experience.

01 / PRIVATE ARCHITECTURE

A companion up front.
Private compute behind her.

Echo receives your request and stays with it. A control node routes specialist work to your private compute pool.

Architecture plan
Your continuous point of contactEcho

Conversation · Presence · Review

Control node / SparkPlanner & Router

Plan tasks · Allocate resources · Track progress

Your private boundary
Private Compute Pool

Research / Code / Video / Review

Memory · Permission-scoped context
Private files and retrieval libraries stay outside tenant workloads
Optional external services

Named provider, task, data and budget.
Only approved context leaves after consent.

Cloud models / Burst GPUs / Hosted 2K

External workloads are separate from personal Memory.

Explore companion interactions and privacy controls

The companion comes first.

Still here, even when the answer takes time.

Echo acknowledges you, keeps the conversation moving, and shows what is happening while deeper work continues.

Presence is real-time; intelligence can take time.

Illustrative interaction · product target

“Can you work through these notes?”

Echo

I’m here. Let me work through that.

One experience, across the proposed range.

Present while thinking

Talk naturally, interrupt, and know when Echo needs more time.

Memory you control

Keep useful context close, with controls to review, delete, and pause memory.

Context with permission

Connect the files, apps, and optional voice or visual inputs you choose.

Work you can follow

See what is running, waiting, finished, or needs your attention.

Outside help, by choice

Approve the destination and context before a task uses a cloud model.

Local first. Your choice, always.

Your context. Your decision.

Personal memory stays local by default. Cloud assistance shares only the context you explicitly approve.

  • Choose which files, apps, microphones, cameras, and screens Echo may access.
  • Review and delete saved memories; pause observation and background work.
  • Inspect the exact proposed cloud payload before sharing.
  • Disable cloud inference and continue with local capabilities.

More capacity. The same boundaries.

When outside help would be useful, Echo explains why, shows what would be shared, and asks first. You can always keep the task local.

  1. 01
    Know who, and why.

    Named provider and model, why outside expertise would help, and the proposed task.

  2. 02
    Review the exact context.

    Messages, excerpts, files, images, and tool context prepared locally. Nothing is pre-uploaded.

  3. 03
    Decide for this task.

    Review retention and training terms, charges or a spending cap, and any unknown completion time. Consent applies only to the disclosed task and payload.

Policy explanation only. These controls do not connect to an AI service or send any context.

The consent contract, in full

Before anything leaves

Bind consent to a task, provider, disclosed payload, and cost limit. Any expansion requires new consent. Declining or not responding never authorizes upload.

Prepare the smallest useful payload locally. No background pre-upload, full memory export, or automatic fallback to a second provider.

If you decline

Continue locally where feasible, queue the work, narrow its scope, or explain that it cannot be completed locally.

If you change your mind

Cancel unsent work and stop future transmissions. Explain that already transmitted data cannot be recalled by Echo.

After outside help

Identify when an answer used external assistance and preserve the task’s sharing record.

Disclose the selected provider’s retention and training terms before consent. Never imply zero retention unless the selected service and configuration support it.

Permissions stay yours

  • Require user-granted access to each data source and integration.
  • Keep delegated work within an approved task scope, schedule, and resource budget.
  • Pause at approval checkpoints for publishing, sending, purchases, destructive changes, or other consequential actions unless explicitly pre-authorized within that scope.
  • Make active work inspectable and cancellable; keep a record of actions and exceptions.

What “local” means

Local means your configured device or private deployment. Remote hosting, backups, telemetry, connected apps, and web requests each need clearly disclosed data boundaries; local inference alone does not make those activities offline.

02 / AGENT SCALING

One point of contact.
A team for the task.

From planning and execution to review, Echo organizes complex work into results you can inspect.

Workflow demo · Services planned
01

One assistant

Spark or a 32–72GB GPU

A planner breaks down requests. Lightweight specialists start on demand; heavy work queues or uses authorized external services.

02

Personal team

96GB or 2× PRO 6000

Planning, execution and review have distinct roles. Two GPUs can separate conversation from heavy video jobs.

03

Private work team

4× PRO 6000

Specialist tasks can run in parallel and create subtasks within depth, concurrency and budget limits.

04

Private production

8× PRO 6000

Keep interaction services resident and separate from burst capacity for multiple workflows and batch work.

05

Agent factory

1× B300 → 8× B300 HGX

Evaluate larger model services and heavy workloads with more memory and high-speed multi-GPU interconnects.

EchoPlannerResearcher / Coder / Video WorkerCriticDelivery
Resident roles stay available. Burst roles release resources when done.

Running: 3 demo tasks · Queued: 2 demo tasks

Demo data: the planner splits research, code and video tasks. Explore the next step; no real agents execute.

Agent creation is a software capability. GPU count is not an agent count. Recursive work needs depth, concurrency, budget and timeout limits, plus cancellation propagation. Concurrency requires measurement.

03 / VIDEO AI

From a single shot
to your private production team.

Echo coordinates storyboards, generation, review and editing to turn an idea into work you can review.

01 / Audio + videoMiniMax H3Local + hosted hybrid
NM integration planned

H3-Base / Context-IR / Regenerate-2K

Text, first/last frames and multimodal references; 4–15 seconds, 24 FPS and native stereo audio.

Open Base defaults to a 768p short edge. The full official 2K workflow includes hosted Context-IR and Regenerate-2K.

Official documentation & examples ↗

Official repository examples. Execution location and parameters depend on each example.
Specifications checked 2026-09-07

02 / Local creationWan 2.2Local workflow candidate
NM integration planned

TI2V-5B

Text-to-video and image-to-video at 720p / 24 FPS.

Configure each variant separately. Do not assume the native audio features of H3.

Official documentation & examples ↗

Official model-card examples, not NM local benchmarks.
Specifications checked 2026-09-07

03 / World modelsCosmos 3Simulation + data generation
NM integration planned

Evaluate by variant

Multimodal world generation and action conditioning; primary generation settings up to 720p / 30 FPS.

For physical AI. Check Edge limitations separately; this is not a real-time avatar service.

Official documentation & examples ↗

Official model reference. No unverified NM demonstration footage.
Specifications checked 2026-09-07

04 / Cloud creationVeo 3.1External cloud service
NM integration planned

Cloud API / service

Native audio, reference controls and enhanced consistency.

Access through cloud services. More local GPUs do not make Veo locally deployable.

Official documentation & examples ↗

Official provider showcase. Output settings and conditions are provider-defined.
Specifications checked 2026-09-07

  1. DirectorScript
  2. StoryboardShots / References
  3. Video WorkersGenerate candidates
  4. CriticSelect / Rework
  5. EditorEdit / Audio
  6. YouReview final cut

A 60-second ad is a multi-shot production goal. Offline video is separate from real-time avatars; character consistency and automated review are workflow goals. Linked examples are from official providers, not footage of an integrated NM service.

04 / AUTONOMOUS COMPUTE ECONOMY

Yours when you need it.
Useful when you do not.

Reserve capacity for yourself. Put permitted idle resources to work, and expand beyond local compute within your budget.

Interactive demo · Renting, settlement and automatic purchases planned

Protect Echo and your personal work

External renting is off by default. Reserve interaction capacity and see personal jobs alongside available headroom.

Owner-first

Stop taking new external jobs first. Drain, preempt or wait according to each contract. A UI switch cannot promise instant recovery.

Zero-trust

External jobs require separate execution environments, networks, storage, credentials and permissions. Clear tenant data afterward and retain audit records.

8× B300 / AllocationDemo data
GPU 01Echo reserveInteraction
GPU 02Echo reserveInteraction
GPU 03Personal agentPersonal use
GPU 04Video jobPersonal use
GPU 05IdleEligible · Off
GPU 06IdleEligible · Off
GPU 07IdleEligible · Off
GPU 08IdleEligible · Off

Reserved 2 GPUs · Personal use 2 GPUs · External use 0 GPUs · Idle 4 GPUs

2 queued · No recovery needed · External budget: $0

Demo data. Resources stay local and nothing is sent externally. Four GPUs represent eligible headroom only.

01

GPU-hours

Offer authorized GPU capacity by the hour

02

Inference / tasks

Deliver inference or video tasks

03

Agent outcomes

Deliver completed agent outcomes

Three commercial stages, not three markets launching at once. After reserving two GPUs, the remaining six are not a complete eight-GPU HGX. No simulated live income is shown.

05 / CAPABILITY LADDER

From your first Echo
to parallel production.

Start with DGX Spark for Echo Core, then explore nine dedicated GPU configurations for generation and multi-GPU production. Size for the workload; memory alone does not determine visual quality.

Showing 10 / 10 reference configurations

01Planned · validation required

NVIDIA DGX Spark

GB10 · Blackwell · ARM64 · Personal desktop

128GB unified memory

Shared CPU + GPU memory

Echo Core · Personal AI entry

Start with a compact private node for everyday assistance, document search, drafting and task planning. Add a dedicated generation node as your workload grows.

128GB unified memory is shared by the CPU, GPU and operating system. Heavy tasks may queue; model fit and application compatibility require validation.

Official specifications ↗

Checked 2026-09-07 · Echo Core deployment requires evaluation

02Planned · validation required

RTX 5090

Blackwell · Local workstation

32GB per GPU

1 GPU × 32GB per GPU

Optimized generation entry

Evaluate local generation with quantization, CPU offload or staged loading.

Heavy jobs share resources with Echo. Full component residency is not guaranteed.

Official specifications ↗

Checked 2026-09-07 · H3 deployment requires evaluation

03Planned · validation required

RTX PRO 5000

Blackwell 48GB · Local workstation

48GB per GPU

1 GPU × 48GB per GPU

H3 Creator Node

More loading headroom than 32GB, designed around an optimized pipeline.

Do not assume every component stays in memory at once. Validate the exact model variant.

Official specifications ↗

Checked 2026-09-07 · H3 deployment requires evaluation

04Planned · validation required

RTX PRO 5000

Blackwell 72GB · Local workstation

72GB per GPU

1 GPU × 72GB per GPU

Fewer memory compromises

Keep more components resident and reduce transfers between system and GPU memory.

Both 48GB and 72GB variants have 1,344 GB/s bandwidth. More capacity does not imply proportionally faster generation.

Official specifications ↗

Checked 2026-09-07 · H3 deployment requires evaluation

05Planned · validation required

RTX PRO 6000

Blackwell Server Edition · Compatible server

96GB per GPU

1 GPU × 96GB per GPU

Dedicated video node

Evaluate an optimized H3 service together with video pre- and post-processing.

Coexistence with a large language model depends on measured peak memory.

Official specifications ↗

Checked 2026-09-07 · H3 deployment requires evaluation

06Planned · validation required

2× RTX PRO 6000

Blackwell Server Edition · Private server

192GB aggregate

2 GPUs × 96GB per GPU

Room for conversation and creation

Reserve GPU 0 for interaction and GPU 1 for video generation.

192GB is aggregate capacity. Split services across cards; it is not one unified memory pool.

Official specifications ↗

Checked 2026-09-07 · H3 deployment requires evaluation

07Planned · validation required

4× RTX PRO 6000

Blackwell Server Edition · Private server

384GB aggregate

4 GPUs × 96GB per GPU

Parallel creative studio

One GPU can serve interaction and orchestration while three generate shot candidates.

384GB aggregate. Shot concurrency and generation speed require workload benchmarks.

Official specifications ↗

Checked 2026-09-07 · H3 deployment requires evaluation

08Planned · validation required

8× RTX PRO 6000

Blackwell Server Edition · Private rack

768GB aggregate

8 GPUs × 96GB per GPU

Personal video production pool

Separate interaction, generation, reference images and post-production to scale workflow throughput.

768GB aggregate. Budget separately for CPU, RAM, storage, power and networking.

Official specifications ↗

Checked 2026-09-07 · H3 deployment requires evaluation

09Planned · validation required

1× B300

Blackwell Ultra · HBM3e · Compatible data-center server

288GB per GPU

1 GPU × 288GB per GPU

Large-memory H3 node

Evaluate BF16 and larger working sets with less aggressive quantization and offload.

A server resource tier, not a desktop GPU replacement. No guarantee that every workflow fits.

Official specifications ↗

Checked 2026-09-07 · H3 deployment requires evaluation

10Planned · validation required

8× B300 HGX

Blackwell Ultra · NVLink / NVSwitch · Data-center cluster

2,304GB aggregate

8 GPUs × 288GB per GPU

Agentic Media Factory

Use high-speed interconnects for multi-GPU work, shot pools, candidate batches and editing.

2,304GB aggregate. Frameworks must explicitly support sharding and parallel execution.

Official specifications ↗

Checked 2026-09-07 · H3 deployment requires evaluation

Aggregate memory ≠ single-GPU memory. These are deployment recommendations, not measured performance tiers. On 32–96GB, feasibility depends on quantization, offload, input size and software version. Hardware changes loading headroom, latency, concurrency and scaling. Published benchmarks need a model version, precision, resolution, duration and test conditions.

06 / INFRASTRUCTURE

Start with your work.
Plan room to grow.

Plan models, interaction capacity, storage, networking, power, cooling and maintenance together. Pricing requires a defined configuration and a current quote.

Configuration planning

Personal node

DGX Spark · Add a generation node when needed

Validate your primary workflow first.

Plan this stage ↗

Private teams

2× / 4× PRO 6000

Reserve resources for interaction and creation.

Plan this stage ↗

Research clusters

1× B300 → 8× B300 HGX

Validate large working sets, interconnects and whole-node scheduling.

Plan this stage ↗
Explore the five Echo product tiers and comparison

The five product tiers describe service plans; the hardware configurations are references, not one-to-one equivalents. Explore the original comparison, download it or create a local inquiry brief.

Find your Echo.

Choose your level of local independence.

Every Echo tier can seek outside expertise with your permission. Higher tiers give more of your work room to happen locally, together, and over time.

Proposed range · final capabilities and availability to be validated.

Compare all five

Echo Core · Guided tasks

Your companion starts here.

Start with Echo at your side: remembering what matters, helping with everyday work, and finding a way forward when a task needs more.

Designed forFor individuals who want private everyday assistance in a compact personal system.

Ways to put it to work

  • Recall a detail

    Find a decision in your local notes and show the supporting passage.

  • Prepare your day

    Draft a brief from approved local information and flag missing context.

  • Work through a request

    Summarize a document, create a draft, and ask before sharing it.

Your context, close to you

Bring your notes, documents, and everyday context into one companion.

Help that follows through

Turn a request into a short, visible sequence of actions.

A path forward

Wait locally, simplify the task, or approve outside expertise.

Local capabilities & boundaries

Proposed local capabilities

  • Voice interaction, interruption, and lightweight visual context when enabled.
  • Personal memory search, document summaries, drafting, and everyday planning.
  • Short tool workflows using approved integrations; background indexing scheduled around conversation.

Guided tasks

Completes short, bounded workflows with checkpoints; proposes extra resources when a task outgrows the local system.

Light background work, with heavier tasks queued.

What to plan for

  • Limited room for large models and simultaneous heavy tasks; long jobs may queue.
  • Shared memory also serves the operating system and CPU; 128 GB is not fully available for model weights.
  • Local quality depends on the chosen models. Some tasks may remain unsolved if cloud assistance is declined.

Outside help, by choice

Expect to consider outside help for difficult reasoning, large workloads, or tight deadlines.

Continue locally where feasible, queue the work, narrow its scope, or explain that it cannot be completed locally.

Reference hardware & deployment

~128 GB shared accelerator-addressable memory

DGX Spark-class personal appliance

Single integrated CPU/GPU system with shared memory.

Compact desktop appliance · 1 reference node

NVIDIA DGX Spark system overview ↗
  • No multi-node cluster required for this reference.
  • Leave memory and scheduling headroom for speech, perception, the OS, and runtime state.
  • Validate the full conversation stack under concurrent indexing load.

Hardware references are illustrative, not certified Echo configurations or performance benchmarks. Installed memory is not fully available to models.

Proposed capabilities; final availability and performance depend on the validated configuration. Reference hardware is illustrative. Cloud assistance is optional and may incur separate charges.

Hardware determines how independent Echo is, not how intelligent Echo is allowed to be.

More compute increases capacity to act independently within a defined scope; it never grants broader permissions.

The right kind of room.

Compare your Echo.

Compare the intended local work envelope. Responsive presence and control over sharing are common to every tier.

Save the comparison
Proposed Echo tier capabilities
Designed forCorePlusProMaxInfinite
Local capabilityAssisted local capabilityExpanded local capabilityAdvanced local capabilityDedicated private computeFlagship private compute
Delegated workGuided tasksRoutine workflowsDelegated projectsOngoing operationsCustom ongoing operations
Background workLight background work, with heavier tasks queued.Routine background workflows alongside conversation.Longer projects and multiple scheduled workflows.Persistent workers across several demanding projects.Broad persistent workloads with custom capacity planning.
Responsive presenceSame design goal on every tierSame design goal on every tierSame design goal on every tierSame design goal on every tierSame design goal on every tier
Default privacyLocal first; explicit consent for cloud inferenceLocal first; explicit consent for cloud inferenceLocal first; explicit consent for cloud inferenceLocal first; explicit consent for cloud inferenceLocal first; explicit consent for cloud inference
Outside assistanceOptional, with task-specific consentOptional, with task-specific consentOptional, with task-specific consentOptional, with task-specific consentOptional, with task-specific consent

Scroll across to compare all five tiers.

Compare reference hardware
Proposed Echo tier capabilities
Reference hardwareCorePlusProMaxInfinite
Reference memory~128 GB shared accelerator-addressable memory~192–384 GB aggregate VRAM~384–768 GB aggregate VRAM~1 TB HBM-class; reference total 1,128 GB~2+ TB HBM-class; reference total 2,304 GB
Reference computeDGX Spark-class personal appliance2–4 × RTX PRO 6000 Blackwell-class GPUs4–8 × RTX PRO 6000 Blackwell-class GPUs8 × H200 in an HGX/DGX-class system8 × B300 in a DGX/HGX-class system

Plus and Pro intentionally overlap at 4 × 96 GB. Assign the tier by a validated workload envelope, including reserved presence capacity, solver workload, concurrency, and supported workflow software. A GPU count alone cannot resolve the overlap.

Aggregate multi-GPU memory is not automatically a single fast memory pool. Fit and performance depend on sharding, quantization, context/KV cache, activations, runtime overhead, and interconnects. Spark’s 128 GB is shared CPU/GPU memory.

NVIDIA DGX Spark system overview ↗
NVIDIA RTX PRO 6000 Blackwell Server Edition ↗
NVIDIA DGX H100/H200 system introduction ↗
NVIDIA DGX B300 system introduction ↗

07 / TRUST & FAQ

Move forward.
Know the boundaries.

Clear answers on capabilities, privacy and resource use.

What is available today?

This page supports hardware browsing, workflow illustrations and allocation demos. Echo model services, video generation, capacity markets, automatic purchasing and settlement are planned. Hardware tiers are design guidance, not benchmarks. A saved receipt confirms an inquiry submission.

Does 2×96GB behave like a single 192GB GPU?

No. Multi-card memory figures are aggregate. PRO 6000 configurations prioritize service separation. Even B300 HGX with NVLink / NVSwitch needs explicit framework support for sharding and parallelism.

Do I need B300 to start with H3?

No. Evaluate quantization, offload or staged-loading paths on 32–96GB configurations. Feasibility depends on the variant, precision, input size and software version. There is no universal memory threshold for a complete workflow.

Is the full H3 2K workflow local?

The official complete pipeline includes hosted Context-IR and Regenerate-2K, so it is presented as a hybrid service. Before execution, disclose the materials sent, destination and costs. Users can choose to keep work local.

Can it generate a 60-second ad in one pass, or a real-time avatar?

A 60-second ad is a multi-shot generation, review and editing goal, not a single 60-second H3 output. Offline video and low-latency avatars use different service chains. Character consistency and automated review do not succeed on every attempt.

Can rented GPUs return to me immediately?

Personal capacity should be reserved. When demand rises, stop accepting external work first, then drain, cancel preemptible work or wait for the contract to expire. Show release timing and alternatives when immediate recovery is unavailable.

Will outside tenants see my private memory?

Private memory, retrieval libraries and files are excluded from the proposed tenant data flow. External workloads require separate execution environments, networking, storage and credentials. Isolation must be verified in the backend; a diagram alone is not a security guarantee.

What revenue or payback can I expect?

There is no real settlement data yet, so this page makes no live-income or payback claims. Future estimates must disclose price, hours, utilization, electricity and other costs. GPU-hours, tasks and finished outcomes are separate commercial stages.

Can I rent a full eight-GPU HGX after reserving two GPUs?

No. Six available GPUs are not a complete eight-GPU node. Whole-node jobs require safely migrating personal services or providing separate reserved capacity within your rules.

What is sent with an inquiry?

Only your entered name, email, intended work, scale and selected configuration are submitted to this website for infrastructure consultation. No personal AI memory or private files are sent. Do not include passwords, keys or sensitive business information. Local brief copying and downloading remain available.

08 / LET’S BUILD YOURS

Plan your private
AI infrastructure.

Tell us what you want Echo to accomplish, and the scale you have in mind.

From research and coding to a private video team. Define your workflow, then choose the hardware that fits.

A saved receipt confirms submission. NM / NM Technology uses your inquiry to respond to your needs. Do not include passwords or sensitive files.

Submitting sends your details to this website. Local briefs remain in this tab. The site does not automatically send email or call AI services.

Your world. Your Echo.

Prepare a brief for a conversation with NM Technology. Nothing is sent from this page; you can copy or download it.

Pricing and availability are not yet specified. Your entries stay in this browser tab.