GEODI 300 - Deployment & Architecture & Performance Design

GEODI 300 - Deployment & Architecture & Performance Design

Purpose

Deploying GEODI begins with understanding how different workloads behave.

This course explains GEODI’s architecture, performance design, and sizing principles. It highlights how Discovery, Classification, and Endpoint workloads require different approaches. You will learn the platform’s core components and system-planning fundamentals, including server capacity, storage, and connectivity.

You will learn how to manage data volume, concurrency, and performance expectations applying the principle: optimise first, then scale.

By the end, you will be able to design efficient, scalable, production-ready GEODI environments with confidence.

Target

General

Prerequisite Trainings

We recommend first reviewing the end-user training. This helps you gain a better understanding of the platform and these courses by focusing on the actual user experience.

Scope

Please use the pages below for detailed content and refer to the website for a higher-level overview.

  1. GEODI-300 Architecture & Performance Design

    1. GEODI-300_System_Requirements

    2. GEODI-300-Single_GEODI-Server_Installation

    3. GEODI-300 Sizing

Questions

Try to answer the Quick Knowledge Check questions at the end of each page to confirm your understanding

Supporting Material

GEODI Installation Requirements

GEODI Topology

GEODI Server Installation

GEODI Download

GEODI Privacy Issues

Data Sources & Integrations

 

 

🚀 Quick Knowledge Check

🎯 GEODI-300 / Architecture & Performance Design

Because each workload has different performance, continuity, and infrastructure requirements. Treating them the same can lead to inefficiency and instability.

It ensures data security, compliance, and reduces network overhead while allowing processing without moving sensitive data.

Server capacity, storage, connectivity, secure access, and operational access form the foundation of any GEODI deployment.

It ensures stability, automatic execution, and uninterrupted operation required for production environments.

Data volume, data type, index strategy, OCR usage, and scan scope all directly influence Discovery performance.

It means improving scan scope, indexing strategy, and data segmentation before increasing hardware resources.

Because it runs continuously in production, handling automation and policies, making downtime a direct operational risk.

Large environments and high endpoint volume require distributed processing to maintain performance and scalability.

They provide module updates, licensing, templates, and support services necessary for system health and continuity.

Separating workloads, planning for continuity, aligning infrastructure with performance needs, and ensuring operational visibility.

 

🎯 GEODI-300 / Server Requirement

Because it enables centralized management and access without requiring installation on user devices.

Because discovery and indexing operations are I/O intensive and require high-speed disk performance.

Because indexing during discovery typically consumes an additional 10–20% of storage.

Because secure communication between server, agents, and clients is required in production environments.

Because it allows module updates, templates, and AI assistant features without exposing the system to full internet access.

Read-only is used for discovery, while write access enables remediation actions.

Because they have different performance characteristics and continuity requirements.

Because large data volumes and limited scan windows may exceed the capacity of a single server.

Because it is CPU-intensive and increases processing time during discovery.

To provide and manage AI infrastructure, including hardware resources, LLM setup, and API credentials.

 

🎯 GEODI-300 / Sizing

Because Discovery is mainly driven by data volume and scan scope, while Classification requires continuous, low-latency performance and high availability.

Classification always requires High Availability. In practice, this means starting with at least two servers. If concurrency, traffic, and performance requirements exceed the capacity of those servers, additional servers should be deployed to increase capacity.

Sampling/QuickScan, index strategy, number of GEODI nodes, and data source type directly impact throughput, storage, latency, and resilience.

Because tuning sampling, index options, and segmentation can significantly reduce workload without increasing infrastructure cost.

Full Text provides complete search capability but requires more storage and time, while Metadata Only is faster and smaller but limits search depth.

It reduces processing time and index size proportionally, but results are limited to the sampled portion of data.

Because it is CPU-intensive and can significantly slow down discovery, so it should be applied selectively or delayed when possible.

To ensure stability, continuous operation, and predictable performance without being affected by Discovery workloads.

Data volume and endpoint count, with a common guideline of one server per ~10 TB or ~1,000 endpoints.

Because distributing endpoint load across multiple servers improves performance, reduces bottlenecks, and enables scalable discovery operations.

 

🎯 GEODI-300 / Single GEODI Server Installation

Because service mode ensures continuous, unattended operation and automatic startup, which is required for production environments.

Because it stores indexes, logs, and projects, and incorrect placement or permissions can directly impact performance and system stability.

Modules cannot be automatically downloaded or updated, making setup and maintenance more complex and manual.

Because it allows secure module updates and template synchronization without exposing the system to unnecessary external access.

Because clients and endpoints require a secure and reachable server address for communication, especially in production environments.

It provides centralized access to configure, monitor, and manage GEODI projects and system settings.

Because GEODI continuously reads and writes data there, and insufficient permissions can cause failures or data inconsistencies.

They provide ready-to-use configurations aligned with regulations and are automatically updated when online.

By allowing SIEM tools to directly read GEODI log files for near real-time monitoring and analysis.

Because GEODI architecture is designed to grow, and early design decisions affect future scalability and high availability options.