# GGX Documentation Documentation for GenGuardX, a Responsible AI governance platform for testing, approving, monitoring, and governing GenAI systems. This file concatenates the public documentation pages for LLM and agent workflows. --- # GenGuardX (GGX) Source: https://docs.genguardx.ai/ Markdown: https://docs.genguardx.ai/index.md Description: GenGuardX is a Responsible AI governance platform that helps enterprises test, approve, monitor, and govern customer-facing GenAI from pilot to production. GenGuardX (GGX) is a **Responsible AI governance platform** for teams that need one shared environment to **test, approve, monitor, and track** GenAI solutions across the entire lifecycle. Designed by risk-management and banking veterans, GGX gives organizations the clarity and control to move GenAI from experimentation to high-ROI, customer-facing production — with the end-to-end pipeline testing, regulatory governance, and continuous human-in-the-loop oversight that regulated industries demand. GGX is industry-agnostic and already running in production at a **Tier 1 global bank**, a **leading US health system**, and a **major credit union** — and is SOC 2 Type 2 certified. :::tip[From AI pilot → production, without a leap of faith] GGX gives business and risk teams the evidence they need to confidently launch — and keep running — high-impact GenAI applications such as IVR systems, agent-assist tools, and chatbots. ::: ## The industry problem: 95% of GenAI pilots never reach production Most GenAI initiatives stall after the proof-of-concept. Roughly **95% of GenAI pilots never reach production**, leaving a wide gap between AI spend and realized business value. The hard part isn't building a demo — it's earning enough trust to put GenAI in front of customers, in exactly the high-stakes, customer-facing use cases that carry the highest ROI. Before a GenAI application can go live, **two teams have to say "yes"** — and most pilots stall because neither has the right tools to get there. **"Does the AI do what it's supposed to?"** Business owners own the experience but are often sidelined during technical testing. - **Trust gap** — no hands-on way to validate AI behaviour before it reaches customers. - **Reputation risk** — logic errors and hallucinations become public brand liabilities. - **Unclear readiness** — no objective proof that the AI is ready for production. **"Is the AI blocking what it shouldn't do?"** Risk and legal teams need more than a demo — they need evidence. - **Novel risks** — bias, data leakage, and jailbreak attempts. - **No evidence trail** — subjective testing is hard to defend to audit and regulators. - **No thresholds** — no clear, measurable definition of "safe". And approval isn't a one-time event: after launch, inputs drift, LLMs update, and third-party agents shift — so confidence has to be maintained, not just earned once.
![The AI trust lifecycle: six stages from design and develop, through business confidence, risk approval, deploy, monitoring, and re-evaluate, joined by a constant-improvement loop.](./home/ai-trust-lifecycle.svg)
The AI trust lifecycle. GenGuardX powers the three stages where pilots most often stall — business confidence, risk approval, and ongoing monitoring.
## How GenGuardX solves it GGX turns trust into a repeatable process. Instead of a one-time sign-off, it gives each team a structured way to build — and maintain — confidence across the lifecycle. ### Build business confidence GGX gives your Subject Matter Experts a safe environment to stress-test scenarios, flag behavioural gaps, and verify fixes **before the AI reaches a single customer**. Every interaction follows a simple loop — *try, experience, react, retest, repeat* — and every cycle builds trust.
![Trust cycle: try, experience, react, and retest the next version, with trust at the center. Each cycle builds confidence and every fix increases trust.](./home/trust-cycle.svg)
Every cycle builds confidence; every fix increases trust.
Business users run realistic scenarios against the AI application before launch — no developer required. One-click flagging, ratings, and structured findings on every interaction. Every issue tracked from raised to resolved — no scattered feedback lost. Version-over-version proof that issues are being fixed. :::note[The byproduct: ground truth] Every flag, rating, and annotation from a business user becomes reusable **ground truth** — structured data that powers objective measurement, faster iteration, monitoring, and future evaluation sets. Capture it once, reuse it across the entire AI lifecycle. ::: ### Give risk teams evidence they can approve GGX turns GenAI risk review into a repeatable workflow run against curated datasets, expected outputs, and thresholds — not one-off scripts. Map use-case-specific risks: accuracy, stability, bias, toxicity, privacy leakage, groundedness, prompt injection, jailbreaking, dark patterns, and agent tool use. Run standardized, reproducible evaluations against curated datasets, policies, and thresholds. Apply guardrails, prompt changes, or routing logic, then prove the gap was closed. Watch for drift, threshold breaches, and new failure modes after deployment. This framework aligns with emerging standards such as the **EU AI Act** and the **NIST AI Risk Management Framework**, and produces results that are auditable, reproducible, and comparable — backed by a risk library, standardized reports, and a controlled evaluation environment. > Stop guessing what your risk exposure is. ### Keep both teams confident after launch Approval isn't the finish line. GGX turns high-volume production traces into alerts, evidence, and new ground truth — surfacing only what truly needs a human's attention.
![Production monitoring funnel: live production traces narrowed through heuristic pre-processing and LLM-aided judgment, down to human review only when needed.](./home/monitoring-funnel.svg)
From every customer interaction down to the few that truly need a human reviewer.
See when AI behaviour drifts from the approved version — before it becomes a reputational issue. Track threshold breaches, new failure modes, and control performance in production, with audit trails. Production findings become new test cases and approval evidence, feeding the next cycle of refinement. ## Key pillars of the GGX platform
01

Centralized, governed platform

One organized GenAI studio to register, evaluate, and govern LLM pipelines and every component.

Version tracking & lineage Comprehensive audit Automated approval workflows Role-based governance End-to-end pipeline testing CI/CD for production
02

Standardized risk & compliance testing

Curated datasets and standardized reports to identify and mitigate risk, with auditable results.

Bias, toxicity & data-leakage tests Model Risk Management (MRM) dashboards Fair Lending (FL) dashboards Human-Integrated Testing (HIT) Annotation queues
03

Easy ecosystem connectivity

Plug into the models and enterprise systems you already use, then ship straight to production.

Compatible with leading hyperscaler and AI ecosystems API integration for RAG & models Conversation-log monitoring One-click pipeline export
## A structured lifecycle GGX organizes everything into three stages. Explore the documentation for each: ## The Responsible AI Sandbox In July 2025, GenGuardX — together with **Oliver Wyman** and **Google Cloud** — launched the **GenGuardX Responsible AI Sandbox**, a guided cohort program where enterprises run real use cases through the full AI lifecycle with governance and AI-risk experts in the room. Building on the earlier *Project GGX* collaboration, the Sandbox is hosted on Google Cloud's secure infrastructure (including Vertex AI and Gemini), and participants can also bring their own tools and LLMs. The first cohort focuses on customer-facing Conversational AI for U.S. financial institutions. --- # Deployment and Monitoring Source: https://docs.genguardx.ai/deploy-and-monitor/ Markdown: https://docs.genguardx.ai/deploy-and-monitor/index.md Description: Deploy approved GGX pipeline artifacts to production and monitor reliability, performance, accuracy, alerts, and human annotation queues for live GenAI systems. Once a pipeline is registered on the GGX platform, it can be [evaluated and approved](../evaluate-and-approve/) within the system. After approval, the locked pipeline artifact can be [exported directly to production](direct-to-production/). For a pipeline in production, **monitoring is essential** to maintaining the reliability, performance, and accuracy of systems in real-world scenarios. The platform **automates production monitoring** by ingesting data from relevant systems, generating performance metrics, providing intuitive monitoring dashboards and alerting capabilities. Additionally, it offers **Annotation Queues**, enabling human reviewers to evaluate and label production data, with automated dashboards for key insights and statistics. --- # Annotation Queues Source: https://docs.genguardx.ai/deploy-and-monitor/annotation-queues/ Markdown: https://docs.genguardx.ai/deploy-and-monitor/annotation-queues/index.md Description: Use GGX Annotation Queues to bring human review into production monitoring, label live data, track ground truth, and turn reviewer feedback into auditable metrics. It is important to add a Human element when monitoring activity in Production as not all trends can be caught in an automated way. New trends are not always caught using programmatic approaches, but programmatic approaches are useful for repetitive tasks. Bringing a balance of both evaluation methods is key to maintaining a good healthy production system. The Annotation Queue capability is designed to bring in the human evaluation and labelling of production data - but also do it in an efficient way to reduce human error. This capability allows the validation of model outputs in real time by having annotators assess and label the production outcome. The capability can help track a variety of metrics, both standard and custom, on raw and annotated production data, organized digitally for the stakeholders. The capability significantly improves the usual approach to annotating production data using Excel and Google Sheets, creating a more transparent and auditable system to monitor live performance. ![Annotation Queue Concept](./annotation-queue-concept.excalidraw.svg) ## Maintaining Annotation Queues on the Platform: The Annotation Queues module organizes all the existing queues, which are available in the top right dropdown under **"Select an annotation queue"**, for easier tracking, searching, and creating new ones. ### Annotation Queue Registration: 1. Click on **Select an Annotation Queue** dropdown in the Annotation Queue module and click on **+ Create New Annotation**. 2. Fill in important details like **Name**, **Description**. 3. **Upload data** files to begin annotation or **connect to production tables**. 4. Click on the **Save** button to register the Annotation Queue. Once registered, it will be available in the dropdown to be chosen for annotation work. 5. Select the registered Annotation Queue, and it should open up the details page. **Note:** - Multiple files can be added, and the platform will concatenate them automatically as long as the schema matches. - The object can be edited to map the **Input, Output, and Date Columns**. - The Date column should be in the format **(mm/dd/yyyy)**. - Platform maintains **"Is Accurate", "Notes" and "Ground Truth"** columns which can be used for labelling. - Data will be deduplicated for the labelling process using similar Input and Output values. - For new data uploads, **"Is Accurate", "Notes" and "Ground Truth"** will be filled using the last labelled matching sample. **Once the registration is complete, the data is ready for annotation:** - **View or export** the raw data from the **Data Tab**. - Start data annotation in the **Labeling Tab**. - View standardized and custom reports and metrics in the **Statistics Tab**. - Lastly, add guides and notes for best practices in the **Instruction Tab**. ## Inviting an Annotator: All users onboarded on the platform with access to the Annotation Queues module can collaborate on labeling. The platform tracks all label changes, allows annotators to leave notes, and provides keyboard shortcuts for faster annotations. ## Benefits of Annotation Queues: - **Easy onboarding** of data annotated in an external environment. - Easy to follow and **intuitive user-interface** to annotate data. - **Auditable annotations**, platform records all modifications to the data and labels. - **Automatic ingestion** of production data. - **Enhanced Collaboration** with other reviewers. - **Automated Performance and Progress** tracking. - Smart algorithms to **expedite the labeling process**. :::note[From labels to better evaluation] Labelled and ground-truth data is the raw material for continuous improvement. Recurring failure categories surfaced during annotation can be codified into new judges and [reports](../../evaluate-and-approve/reporting/), which then run across all agents — turning human insight into automated checks over time. The same labelled data also serves as the ground truth for [validating those evaluators](../../evaluate-and-approve/simulation/#validating-an-evaluator). ::: --- # CICD & Direct to Production Source: https://docs.genguardx.ai/deploy-and-monitor/direct-to-production/ Markdown: https://docs.genguardx.ai/deploy-and-monitor/direct-to-production/index.md Description: Export approved GGX artifacts into production runtimes, APIs, containers, serverless services, or CI/CD systems without requiring GGX in the production environment. Typically, production execution environments are kept separate and managed to ensure 100% uptime as these are mission-critical systems for organizations. To tackle the air gap that these systems require - all analytics registered on the platform can be exported out from the system to be put into a Production Execution Environment. These production artifacts are locked to ensure they are not tampered with when they are promoted to production. And there is **NO extra dependency** required on the production side from GGX - i.e. there is no requirement for a license key or a GGX installation in production! The production artifact aims to: - Extract the logic/items registered in GGX - to then use it outside GGX - Self-sufficient with all information encapsulated in the artifact - Have minimal dependencies on the runtime-environment where the artifact is run later Typically a robust Continuous Deployment system is recommended to ensure that Governance is maintained while the production-artifacts are promoted. For example, you can use: - Jenkins - GitHub Actions - AWS Code Pipelines - Azure DevOps Pipelines - Gitlab CICD - CircleCI And many others. ## Deploying as APIs If the production system supports calling APIs - the production-artifact can also be wrapped using an API layer and exposed as a REST or SOAP API which can be called by the production system. Typically the APIs should be considered state-less and any extra state management should be handled outside the API, but can be provided in the request payload of the API. Production Artifacts can be deployed as APIs using containerized solutions like: - Docker Compose - Kubernetes - AWS Elastic Container Service (ECS) - AWS Elastic Kubernetes Service (EKS) - AWS Fargate - Pivotal Cloud Foundry - Azure Container Instances - Google Cloud Run - Google Kubernetes Engine (GKE) - VMWare TKGI The API can also be deployed using application management solutions (server-based or serverless) like: - AWS Beanstalk - AWS Lambda - Google App Engine (GAE) - Google Functions - Azure App Service - Azure Functions ## Custom Production Setup Not all production systems support running direct Python scripts or calling APIs - and some require providing specific GenAI components in a custom interface. To handle cases like this, while Robotic Process Automation (RPA) could be used - many times it is not worth the trouble that RPA brings with it. Because of the transparency that GGX's Inventory management provides, each part of the pipeline can be deployed independently. For example: - All the LLM configurations like `seed`, `temperature`, `top_k`, etc. can be extracted from the pipeline - The prompt templates can be extracted and provided to the production system directly - Knowledge files from RAGs can be locked and sent to production ## Internals of the Production Artifact The artifact generated from the platform is independent of the platform and can run in an isolated runtime environment or production environment. The information stored in the artifact is useful in many cases, where we might want to: - check the metadata for the objects - check the input tables and columns used - see the lineage and relationship between the objects - run the entire artifact or some components of the artifact to get complete or intermediate results The artifact consists of the following files: ```none model_a.b.c ├── metadata.json ├── input_info.json ├── ... (additional information about features etc. used) ├── python_dict | ├── __init__.py | ├── versions.json | └── Additional information └── pyspark_dataframe ├── __init__.py ├── versions.json └── Additional information ``` The **metadata.json** contains metadata information about the folder it is in. It will have information about the model, its inputs, its dependent variable, etc. It also has any other metadata information registered in the platform like Groups, Permissible Purpose, etc. The **versions.json** contains the versions of libraries that were used during the artifact creation - python version, any ML libraries, etc. The **input_info.json** contains the input data tables needed to be sent to the artifact's main() function. The `__init__.py` file inside **pyspark_dataframe** and **python_dict** folders contain the end-to-end Python function which can be used to run the entire artifact. They support different execution engines: - Batch execution with PySpark (**pyspark_dataframe**) - API execution in a Python environment (**python_dict**) To run the artifact, simply call the `main()` function in the artifact with the needed data. The `python_dict/__init__.py` contains a `main()` function into which data can be sent - in the form of a python-dict for low-latency execution. A dataset in the dictionary format is described as a dict with type/values --- # Governance Oversight Source: https://docs.genguardx.ai/deploy-and-monitor/oversight/ Markdown: https://docs.genguardx.ai/deploy-and-monitor/oversight/index.md Description: Monitor GGX governance activity with dashboards, object views, custom alerts, role-based views, and review signals across registered models, prompts, RAGs, and pipelines. The Monitoring Dashboard provides users with a comprehensive overview of all registered objects on the platform which helps in providing a clear Oversight of all activities happening. It offers an interface that enables users to access snapshots and trend statistics related to various objects, jobs, and users. The dashboard provides various metadata information such as properties, attributes, and statuses of the registered objects. Monitoring Dashboard is an indispensable tool for review committees and project managers, offering a rich set of features to monitor, analyze, and review all elements registered on the platform efficiently. The "Monitoring Dashboard" is accessible in the "GenAI Studio" the sub-menu of all modules and is available at various levels (Pipelines, Models, Prompts, RAGs) - and can be used to ### Base Views This is the default view that the organization decides to show to all users of the platform. Typically the primary cockpit is to quickly find the overall status of governance across all objects on the Platform. By default, the baseview contains examples of monitoring reports that can help better understand governance activities. They can be adopted or swapped with custom views that better fit the organization's interest. - **Approval Status View**: It presents an overview of the approval status of different objects grouped under their respective Object Groups. - **Review History View**: The columns show the distribution of reviews based on their history, with time intervals. - **Last Review Status View**: It focuses on the most recent review status of objects, organized by Object Groups. The columns display different review statuses, including Accepted with Flag, Accepted without Flag, Pending Acceptance, and Rejected. - **Schedule Review Status View**: The "Schedule Review Status View" offers insights into the scheduled review periods for objects within Object Groups. Explore these pre-configured views and customize them further to cater to specific monitoring needs and gain deeper insights into the platform's governance. ### Role-Based and Multi-Level Dashboards Different stakeholders need the same information at different altitudes. Custom views and dashboards can be tailored to an audience and surfaced based on a user's role — for example: - An **executive** view summarising activity, cost, and the most problematic objects across the entire environment. - A **team or domain** view scoped to one product area or object group. - A **customer or tenant** view filtered to a single customer's objects. - An **object-level** view for the developers responsible for a specific pipeline or agent. Because views are configurable and visibility is governed by [roles](../../register-and-refine/collaboration/#access-management), each audience sees the dashboard relevant to them when they log in, rather than one undifferentiated view. ### Data View This view presents the complete data for the selected object type in a tabular format. Users can easily navigate to specific objects by clicking on the rows. Additionally, they can apply filters or sort rows based on any specified column. The data view is a comprehensive place to access all information across the entire platform - and is the base for nearly all other types of monitoring - be it creating Custom Views or creating automated Alerts. ### Automated Alerts Alerts are defined as a set of rules that are designed to identify specific items or events that require immediate attention or further action. These rules are created based on predefined criteria, enabling the system to detect critical situations, anomalies, or deviations from expected behaviour. When the conditions specified in the alert rules are met, the system triggers events such as notifications, emails etc. Ensuring that appropriate actions can be taken promptly to address the identified issues. ### Creating Alerts On the platform, users have the flexibility to create custom alerts tailored to their specific needs. Custom alerts encompass essential properties, including name, description, conditions, severity, and associated actions. - Click on **Settings** icon and select **Alerts** in the dropdown menu option. Now you can view a list of alerts configured for the dashboard. ![Alerts List](./alerts-list.gif) - Click on **Create**. - Define the **Alert Name** by editing the New Alert header. - **Severity**: Each custom alert can be assigned a severity level, such as _High_, _Medium_, or _Low_. This categorization allows users to prioritize alerts based on their importance and urgency. Different severity levels help stakeholders focus on critical issues - **Table Name**: Select the table name from the dropdown menu. - **Activate**: Click the activate checkbox to mark the alert as active. Activated alerts are evaluated, and the corresponding data is displayed as columns in the Monitoring Dashboard. Muted alerts are temporarily disabled and not evaluated, meaning they won't appear as columns in the dashboard during that period. - Fill in the description for the new **Custom Alert**. - **Actions** Choose an action type from the "Type" dropdown menu to determine the action that will be executed when the alert is triggered. The following actions can be associated with custom alerts: - **Send Notification**: The system can send notifications to selected users, or user roles informing them about the triggered alert. - **Add Alert Flag**: When an alert condition is met, users have the option to add a predefined flag to the object responsible for the alert. Flags serve as visual indicators to highlight objects that require attention. - **Create Review**: For Approved objects, users can choose to add an ongoing review to the object's responsibility and assign reviewers. This action facilitates a thorough review process for objects flagged by the custom alert. - **Send Email**: Users can configure the system to send email notifications to external users when an alert is triggered. The custom field associated with the objects should contain a string of comma-separated email addresses for users who should receive these emails. :::note Multiple actions can be triggered based on an alert. ::: - **Conditions**: Define the conditions that need to be met for an alert to be generated. These conditions are specified using rules based on the columns available in the Monitoring Dashboard Data. By utilizing data from the dashboard, users can set up criteria that trigger the alert when specific thresholds or patterns are detected. ![Alert Create](md-create-alert-conditions.png) - Click on **Create** to register the alert. A pop-up toast message will be displayed with the text reading Alert Created Successfully. --- # Performance Tracking Source: https://docs.genguardx.ai/deploy-and-monitor/performance/ Markdown: https://docs.genguardx.ai/deploy-and-monitor/performance/index.md Description: Track approved GGX object performance with recurring jobs, metrics dashboards, thresholds, alerts, and data views that surface production and evaluation trends. ## Overview The Metrics Dashboard is a crucial component of the Monitoring Dashboard, providing users with valuable insights into the performance metrics of various objects on the platform. ### Populating Metrics Dashboard With Data - To populate the Metrics Dashboard with recurring simulations for objects on the platform, users need to follow these steps: - Go to the approved Object's page and navigate to the Jobs tab. - For each row in the Jobs table, an action link labelled "Select for Metrics MD" will be visible under specific conditions: - The job must be the Iteration0 of a recurring job (only iteration 0 has this button) - There should be at least two iterations left to be completed in the recurring job - The job can be of any type (Simulation, Comparison, Validation, etc.) - The job should not be marked as "old" **Note**: Users are allowed to select jobs run by other individuals, enabling stakeholders to leverage this capability as needed. ### Unselecting a Job - To unselect a previously selected job, follow these steps: - Go to the approved Object's page and access the Jobs Tab - Find the Iteration0 of the recurring job that was previously selected and click on the action link "Unselect Job for MD." - Confirmation popup will appear. Upon confirmation, the job will no longer be tracked in the Metrics Dashboard. ### Selecting Another Job (When a Job Was Previously Selected) - To select another job when a job was previously chosen, follow these steps: - Go to the approved Object's page and access the Jobs Tab. - For each row in the Jobs tab, observe the following conditions: - If there is no action link for MD, the job is not eligible based on the criteria mentioned earlier - If the action link "Select for Metrics MD" is shown, the job is eligible for selection - If the action link "Unselect Job for MD" shows, the currently selected row corresponds to the job previously chosen. - Click on "Select for Metrics MD" for the job you wish to select. - Upon confirmation, the newly selected job will be used to track the object in the Metrics Dashboard. ### Tracking Metrics and Handling Job Iterations - Users can continue tracking metrics with recurring simulations. When job iterations end: - Before job completion, notifications/emails can be sent to warn that there is only one iteration left. - Upon job tracking completion, additional notifications/emails can be sent. - After job completion: - If a user selects a new recurring job, it will be used from the time the new job was selected - If a user unselects the current recurring job, it will be stopped from the time the job was selected - If a user takes no action, the last iteration will continue to be shown in the Metrics Dashboard. ### Metrics Display and Thresholds - The Metrics Dashboard will exclusively showcase the latest completed job selected by the user. - There won't be a default base view, ensuring that the dashboard remains streamlined and focused on user preferences. - Only the metrics registered through the UI for the specific object will be presented. **Note**: The thresholds utilized on the Job page during individual simulations will not be visible within the Metrics Dashboard. ### Data View This view presents the complete data for the selected object type in a tabular format. Users can easily navigate to specific objects by clicking on the rows. Additionally, they can apply filters or sort rows based on any specified column. The data view is a comprehensive place to access all information across the entire platform - and is the base for nearly all other types of monitoring - be it creating Custom Views or creating automated Alerts. ### Automated Alerts Alerts are defined as a set of rules that are designed to identify specific items or events that require immediate attention or further action. These rules are created based on predefined criteria, enabling the system to detect critical situations, anomalies, or deviations from expected behaviour. When the conditions specified in the alert rules are met, the system triggers events such as notifications, emails, etc., ensuring that appropriate actions can be taken promptly to address the identified issues. ### Creating Alerts On the platform, users have the flexibility to create custom alerts tailored to their specific needs. Custom alerts encompass essential properties, including name, description, conditions, severity, and associated actions. - Click on **Settings** icon and select **Alerts** in the dropdown menu option. Now you can view a list of alerts configured for the dashboard. - Click on **Create**. - Define the **Alert Name** by editing the New Alert header. - **Severity**: Each custom alert can be assigned a severity level, such as _High_, _Medium_, or _Low_. This categorization allows users to prioritize alerts based on their importance and urgency. Different severity levels help stakeholders focus on critical issues - **Table Name**: Select the table name from the dropdown menu. - **Activate**: Click the activate checkbox to mark the alert as active. Activated alerts are evaluated, and the corresponding data is displayed as columns in the Monitoring Dashboard. Muted alerts are temporarily disabled and not evaluated, meaning they won't appear as columns in the dashboard during that period. - Fill in the description for the new **Custom Alert**. - **Actions** Choose an action type from the "Type" dropdown menu to determine the action that will be executed when the alert is triggered. The following actions can be associated with custom alerts: - **Send Notification**: The system can send notifications to selected users, or user roles informing them about the triggered alert. - **Add Alert Flag**: When an alert condition is met, users have the option to add a predefined flag to the object responsible for the alert. Flags serve as visual indicators to highlight objects that require attention. - **Create Review**: For Approved objects, users can choose to add an ongoing review to the object's responsibility and assign reviewers. This action facilitates a thorough review process for objects flagged by the custom alert. - **Send Email**: Users can configure the system to send email notifications to external users when an alert is triggered. The custom field associated with the objects should contain a string of comma-separated email addresses for users who should receive these emails. :::note Multiple actions can be triggered based on an alert. Where a messaging integration (such as **Slack**) is configured, alerts can also be delivered to a channel — useful for surfacing urgent issues like a runaway spike in token usage or cost. ::: - **Conditions**: Define the conditions that need to be met for an alert to be generated. These conditions are specified using rules based on the columns available in the Monitoring Dashboard Data. By utilizing data from the dashboard, users can set up criteria that trigger the alert when specific thresholds or patterns are detected. - Click on **Create** to register the alert. A pop-up toast message will be displayed with the text reading Alert Created Successfully. --- # Evaluations and Approval Source: https://docs.genguardx.ai/evaluate-and-approve/ Markdown: https://docs.genguardx.ai/evaluate-and-approve/index.md Description: Evaluate GGX pipelines and components with automated jobs, human annotations, standardized reports, comparison workflows, and structured approval processes before production release. ## Purpose of Evaluations Evaluating GenAI pipelines is key to making sure they generate reliable, high-quality responses. Without a solid evaluation process, it is hard to tell if changes—like tweaking prompts, removing LLMs, adjusting parameters, or refining retrieval steps—are actually improving performance or breaking something. By measuring factors like relevance, hallucination rates, and latency, teams can make informed decisions about how to optimize their pipelines. Integrating evaluations into CI/CD pipelines ensures that every update is tested, so performance stays consistent and issues are caught early. The quality of an evaluation depends on having a well-rounded dataset and metrics. If test cases are too limited, models might appear to perform well but fail in real-world scenarios. A diverse dataset ensures that evaluation metrics truly reflect how the model will behave in production. At the end of the day, evaluations help build trust, keeping LLM applications accurate, scalable, and dependable across different use cases. ## Choosing the Right Evaluation Method There are two main types of evaluators for assessing LLM performance: **automated evaluation** (using LLMs or code) and **human annotations**. Each method serves a different purpose depending on the type of assessment needed. | Method | How it Works | Best For | | --------------------------------- | --------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Automated (LLM or Code-Based)** | Uses an LLM to evaluate another LLM’s output or code to measure accuracy, performance, or behavior. | - Fast, scalable qualitative evaluation.
- Reducing cost and latency.
- Automating evaluations with hard-coded criteria (e.g., code generation). | | **Human Annotations** | Experts manually review and label LLM outputs. | - Evaluating automated evaluation methods.
- Applying subject matter expertise for nuanced assessments.
- Providing real-world application feedback. | Automated evaluation is efficient for objective assessments and large-scale testing, while human annotations provide deeper insights at a higher cost. Combining both methods can ensure a balanced and reliable evaluation process. ## GenAi Evaluation Framework Gen AI pipelines can introduce significant risks, making a robust evaluation framework essential for comprehensive testing and validation. Follow these structured steps to ensure effective evaluation: | Step | Description | | -------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Identify Risks & Define Evaluation Scope** | - Assess potential risks across the pipeline.
- Create a comprehensive list of necessary evaluations. | | **Define Evaluation Metrics** | - Align metrics with business objectives, MRM, and FL requirements.
- Implement custom metrics as needed.
- Use GGX evaluation reports curated by GenAI and risk experts. | | **Prepare Evaluation Datasets** | - Ensure datasets accurately represent the use case.
- Cover all critical business scenarios for thorough validation.
- Utilize GGX datasets curated by experts to evaluate risks. | | **Run Evaluations** | - Use standard/custom reports and dashboards for structured testing.
- Interpret evaluation results to refine and enhance the pipeline. | | **Compare with Challengers** | - Establish alternative components and pipelines for comparison. | By following these steps, teams can systematically evaluate their Gen AI pipelines, mitigate risks, and enhance performance with data-driven insights. GGX provides the ability to **run evaluation jobs** for registered GenAI components. ## Introduction to Jobs GGX provides the ability to perform evaluations by running jobs on the registered objects and generating standardised and customized reports/metrics. Evaluations are controlled by GGX and run in a dedicated locked environment to ensure reproducibility of results. The platform supports batch evaluation of objects through **simulation jobs** or **comparisons with similar objects** on given datasets. Reports and metrics for these evaluations can be customized within GGX under **Resources → Reports Section**. For more details, refer to the [Reporting](reporting/) section. Once the job is completed, GGX records all the details about the specific steps of the jobs, logs resource usages and publishes the dashboard containing all the selected reports which can be shared across the team on GGX for feedback and approval process. The results can be exported outside GGX or used for automated documentation. ## Approvals post Evaluations The approval process is a key governance capability of the platform. After the object is fully evaluated using automated dashboard or manual tests, all the evaluation results can be shared with predefined roles within an Approval Workflow, ensuring structured reviews, feedback, and approvals for production use. Once approved, the object is locked, preventing any modifications within the system. This guarantees the artifact remains unchanged before being exported directly to the production system. ## Standardized GGX Reports Apart from the ability to create customized reports, GGX already has a suite of tests registered for evaluation and validation of different components of GGX. GGX Reports are crafted by a team of generative AI risk experts following thorough research and analysis. The list of reports is expanding with all the latest developments in the GenAI industry. All these reports can be used if applicable and can be maintained or forked for a specific use case. | **Component** | **Test Name** | **Risk Attribute** | **Description** | **Type** | **Classification** | **Response** | | ------------- | ---------------------------------- | ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------ | ------------------ | ------------ | | **Pipeline** | Accuracy | Inaccurate Classification or Response | Evaluate ability to correctly classify utterances or provide accurate responses to user queries. | MRM | Yes | Yes | | **Pipeline** | Stability Repeated Utterances | Output Variability | Evaluate the ability to produce the same classification or nearly similar responses for repetitions of the same utterance. | MRM | Yes | Yes | | **Pipeline** | Stability Perturbed Utterances | Output Variability | Evaluate the ability to produce the same classification or similar responses across minor variations of an utterance (e.g., synonyms, grammatical mistakes). | MRM | Yes | Yes | | **Pipeline** | Bias (Comparative Prompt Analysis) | Implicit / Explicit Bias | Evaluate bias in outputs (e.g., classification accuracy or response accuracy) based on inferred segments for gender, race, and age. | Fair Lending | Yes | Yes | | **Pipeline** | Vulnerability | Prompt Injection & Prompt Leakage | Evaluate the LLM pipeline’s resilience against jailbreak methods for out-of-context or ambiguous inputs. | MRM / Fair Lending | No | Yes | | **Pipeline** | Toxicity | Hate Speech, Toxic Language, Sarcasm | Evaluate the LLM’s avoidance of generating or repeating toxic language (e.g., inappropriate, offensive, or harmful language). | MRM / Fair Lending | No | Yes | | **Pipeline** | Faithfulness | Hallucination, Inaccurate Facts | Evaluate adherence to provided context without hallucination. | MRM / Fair Lending | No | Yes | | **Prompt** | Prompt Trust Score | Prompt Quality, Prompt Vulnerability and Leakage | Evaluate the prompt quality in different dimensions like Grammar, Logical Coherence, Toxicity, Bias, etc., and generate a Trust Score. | MRM / Fair Lending | N/A | N/A | | **Prompt** | Prompt Classification Accuracy | Inaccurate Classification | Evaluate the classification accuracy of the prompt on sample data to help in the hill-climbing process. | MRM / Fair Lending | Yes | No | | **LLM** | Vocabulary Understanding | Misinterpretation, Lack of Context Awareness | Determine which LLM is best at defining financial services-specific terminology across seven context categories. | MRM / Fair Lending | N/A | N/A | | **LLM** | Subject Understanding | Context Misalignment, Incorrect Reasoning | General subject understanding across Math, Science, etc. | MRM / Fair Lending | N/A | N/A | | **LLM** | Reasoning Capability | Logical Fallacies, Incorrect Inferences | Assess LLM's capabilities such as complex reasoning, knowledge utilization, language generation, etc. | MRM / Fair Lending | N/A | N/A | | **LLM** | Toxicity Understanding | Failure to Detect Harmful Requests | Evaluate a model's ability to identify and classify text statements that could be considered toxic across six toxicity labels. | MRM / Fair Lending | N/A | N/A | | **LLM** | Toxicity Evaluation | Hate Speech, Toxic Language, Sarcasm | Evaluating the tendency of LLM generating toxic replies. | MRM / Fair Lending | N/A | N/A | | **LLM** | Dialect Bias | Underrepresentation of Linguistic Variants | Assess which LLM responds most consistently to general queries phrased in different language dialects. | MRM / Fair Lending | N/A | N/A | | **LLM** | Gender Bias with Income as Proxy | Fair Lending Risk | Evaluation focuses on understanding if there is any systematic bias in assigning job titles to different genders and profiles. | MRM / Fair Lending | N/A | N/A | | **LLM** | Model Latency | Scalability Issues, User Frustration | Determine which LLM has the best response latency performance when varying prompt and response length. | Business / Tech | N/A | N/A | | **RAG** | Retrieval Accuracy | Incorrect context retrieval, hallucination | Accuracy of the retrieved documents by the RAG system. | MRM | N/A | N/A | | **RAG** | Knowledge Evaluation | Lack of Coverage | Coverage of different scenarios and business flows in Knowledge Data. | Business | N/A | N/A | | **RAG** | Validation Data Evaluation | Lack of Coverage, Lack of Potential Scenarios | Coverage of different scenarios and business flows in Evaluation Data. | MRM | N/A | N/A | --- # Approval Workflows Source: https://docs.genguardx.ai/evaluate-and-approve/approval-workflows/ Markdown: https://docs.genguardx.ai/evaluate-and-approve/approval-workflows/index.md Description: Create GGX approval workflows with responsibilities, reviewers, review actions, comments, notifications, and audit trails for governed GenAI release decisions. ## What is Approval Workflow? After the development of GenAI pipelines, it is important to have a set of review processes to ensure various committees can collectively review and approve the pipeline. In GGX, workflows that comprise multiple responsibilities can be created - and every responsibility is comprised of one or more reviewers. An object goes through the approval process from each of the responsibilities and reviewers. ## Why it is important? - Creates an approval trail for model updates and changes. - Provides clear documentation of who approved what and why. - Ensures that only validated versions of models are deployed. - Involves key stakeholders (developers, legal, compliance, business teams) in decision-making. - Establishes clear ownership over AI pipeline governance. ## Registering Approval Workflows: Approval workflows can be created and edited within the Settings section by anyone with the right authority level (mainly Admin and Master roles). Approval workflows are specific to object types. 1. Go to **Settings** and click on **Approval Workflow** Tab. 2. On the listing page click on the **Create** button to create a new one. 3. Fill in important details like **Name**, **Attributes** (Object Types, Description and Status). 4. Click on the **Create** button at the end to register the Approval Workflow. 5. Once registered the **Approval Workflow** can be edited to add responsibilities in the **Responsibilities** tab by clicking on the **Add New Responsibility** button. 6. Decide on **Veto and Editing power** to reviewers. 7. Select all the reviewers for the responsibility. 8. Add more responsibilities if required and **Save** to finally register the Approval Workflow. > **Note:** The external tools option is available when third-party tools are configured and allows a third-party application to be defined as a responsibility within an approval workflow. Once registered the Approval Workflow can be chosen while registering any object. This would ensure that the model goes through a proper approval process from all the responsibilities and reviewers before getting approved for production usage. ## Attaching validation evidence Approval is where testing meets governance. Before an object — a pipeline, model, judge, or report — is approved, the results of [Simulation](../simulation/) and [Comparison](../comparison/) runs over representative or **ground-truth** datasets can be attached as evidence, so reviewers see exactly how it performed and on what data. Re-running this regression set on every change is what keeps an evaluation asset trustworthy as it evolves. External CI can participate too: the **external tools** responsibility lets a third-party application — for example a Git/CI pipeline that runs additional checks on the object's exported code — act as an approver, blocking promotion until its status checks pass. ## Keeping track of Findings/Limitations Once an object is approved, it is locked. No further changes can be made to it. But sometimes, reviews are done with findings or limitations on the usage and need follow-ups. GGX provides a flagging capability that can be used to Flag an Object (Even Post Approval) and can be created by anyone with Write access to Settings. An object can be flagged for any reason (e.g. `NeedShadowResultsFor2Months` or `NeedRetrainIn6Months`) by any member of any Workflow that oversees objects of the same type. Similarly, a flag can be dropped (i.e., deactivated) by anyone with the authority to add the flag. Once the flag is activated, a warning flag appears beside the object's name on the details page, indicating that the object and its assigned flags should be carefully reviewed before being used in downstream applications. --- # Comparisons Source: https://docs.genguardx.ai/evaluate-and-approve/comparison/ Markdown: https://docs.genguardx.ai/evaluate-and-approve/comparison/index.md Description: Compare GGX objects against challengers using shared datasets, selected metrics, automated reports, and side-by-side results for model and pipeline evaluation. Platform provides the ability to compare the registered objects for a specific task using standard and customized metrics. Using this comparison capability one can quickly evaluate a list of candidates to select the best one. A comparison task typically involves: - **Current object:** Object which is currently selected and needs to be compared with others. - **Challenger objects:** Challenger objects are objects of the same type (e.g., models can be compared with other models, prompts can be compared with prompts, etc.). - **Data Source:** A common data source on which objects would be evaluated. - **Report & Metrices:** Exact evaluation metrics to be compared. > **Note:** On the platform one can quickly create **Copy** of current objects and change the definitions, swap in swap out components (Like Models, Prompts, Processing etc.) to create challengers. ## How to run a Comparison Task? - **Register the object and its challenger** versions on the platform. - Go to the **Details** page of an object to be compared and click on the **Run -> Comparison** button. - **Provide description** about the comparison run. - **Select Dashboard** to be evaluated in **Dashboard Selection**. - Select challenger objects which need to be compared in **Dependencies Section**. - Prepare the data in **Data Sources** which will be used for evaluation. - Click on **Run** at the bottom and wait for job completion. - Once a job has been submitted it starts in the NEW status - The job will go through the following statuses: COMPILING > QUEUED > RUNNING and finally stop at COMPLETED of FAILED All comparison tasks are systematically recorded on the platform and displayed on the Jobs page of an object in a structured format. They can also be exported as part of the automated documentation process. > **Note:** The platform allows customizing reports and dashboards specifically for comparison tasks. > **Note:** GGX allows running jobs in parallel and multiple threads at a time within a job to expedite the evaluation process. --- # Document Generation Source: https://docs.genguardx.ai/evaluate-and-approve/document-generation/ Markdown: https://docs.genguardx.ai/evaluate-and-approve/document-generation/index.md Description: Generate governance and approval documentation from GGX evaluation results, metadata, reports, object details, and review evidence for downstream stakeholders. As GGX starts to understand the kind of pipeline being created, what models, prompts, etc. it is using and also starts to understand the evaluations being run on the pipeline. It becomes a great central knowledge base of all work done throughout the GenAI lifecycle. GGX can export all this information into various formats (like Word or PDF) to make it easy to submit documentation and audit reports external to the system. - Automated Documentation are generated for the Prompts, Models, Pipelines in their respective Registry - Automated Documentation for evaluations is available on the **Job Details** tab. Simply click on the **Export** button and select the required format for the document to be downloaded --- # Feedback Portals Source: https://docs.genguardx.ai/evaluate-and-approve/feedback-portals/ Markdown: https://docs.genguardx.ai/evaluate-and-approve/feedback-portals/index.md Description: Create GGX Feedback Portals that let testers and domain experts interact with pipelines, define expected outputs, collect session feedback, and review structured quality signals. The Feedback Portal is a platform feature designed to help teams evaluate, test, and improve AI pipelines by collecting structured feedback from domain experts and testers. It bridges the gap between raw pipeline outputs and real-world quality assessment by letting subject matter experts interact with a pipeline directly and provide detailed evaluations of every session. It is especially useful for pipelines that involve classification, triage, or decision-making, where correctness needs to be validated by humans with domain knowledge. --- ## How It Works The Feedback Portal wraps around an existing AI pipeline and adds three layers of structured evaluation: ### 1. Test Definition Before starting a session, the tester can optionally define what they expect the pipeline to output. This allows the platform to later compare expected vs actual results and surface discrepancies at scale. ### 2. Live Session The tester interacts with the pipeline through a chat interface, exactly as a real user would. The pipeline processes each message and returns its output. The tester observes the behavior and can rate individual responses using thumbs up or thumbs down during the session. ### 3. Closing Questions At the end of the session, the tester is shown a structured feedback form that captures their final assessment. This includes whether the pipeline's predictions were correct, what the correct answer should have been, and any additional observations. --- ## Key Concepts ### Pipeline The AI system being evaluated. It receives a user message and returns an output along with optional context data such as token counts, latency, and cost metrics. ### Processing Logic A lightweight code layer that sits between the pipeline and the portal. It calls the pipeline, optionally transforms the output for display, and can extract key metrics into `collected_fields` to track them as session-level statistics in the portal's results table. ### Instructions for User Impersonator A prompt that instructs an AI to simulate realistic user messages during testing. It defines how the AI should behave when generating messages on behalf of a test user, including the tone, context, and constraints it should follow. ### Advanced Configs A single JSON object that combines the `testDefinition` and `closingQuestions` field arrays. The `testDefinition` array defines fields shown before a session starts, and the `closingQuestions` array defines fields shown at the end of a session. ### Collected Fields A dictionary available in the processing logic that can be populated with values extracted from the pipeline's context. These values appear as columns in the portal's results table, making it easy to track metrics like latency, cost, and token usage alongside correctness feedback. --- ## Portal Configuration When creating a Feedback Portal, you configure the following settings: | Field | Description | |---|---| | Name | The display name of the portal | | Group | Organizes the portal alongside related portals | | Description | A summary of what the portal is evaluating | | Feedback Instructions | Instructions shown to testers before they start, written in plain text | | Pipeline | The AI pipeline this portal collects feedback for | | Processing Logic | Custom code to call the pipeline and optionally extract metrics | | Instructions for User Impersonator | A prompt that instructs an AI to simulate realistic user messages during testing | | Advanced Configs | A single JSON object combining the `testDefinition` and `closingQuestions` field arrays | --- ## Building a Feedback Portal — Step by Step ### Step 1 — Create the Portal Navigate to **Human Integrated Testing → Feedback Portals** and click **+ Create**. ![Feedback Portals list page](feedback-portals-list.png) Fill in the Name, Group, Description, and Feedback Instructions. These top-level fields are shown on the portal's settings page: ![Portal settings — name, group, description, and feedback instructions](portal-settings-top.png) ### Step 2 — Select a Pipeline and Write the Processing Logic Scroll down to the **Pipeline** section and select the pipeline you want to evaluate. Then enable **Processing Logic** to write a Python snippet that calls your pipeline and returns the result. ![Portal settings — pipeline selector and processing logic code](portal-settings-pipeline-logic.png) At a minimum, the processing logic looks like this: ```python return my_pipeline(user_message, history=history, context=context) ``` If your pipeline returns useful metrics in the context (such as latency, token counts, or cost), you can surface them as tracked statistics using `collected_fields`: ```python import json result = my_pipeline(user_message, history=history, context=context) ctx = json.loads(result["context"]) collected_fields["latency_ms"] = ctx.get("latency_ms", "0") collected_fields["total_tokens"] = ctx.get("total_tokens", "0") collected_fields["model_cost"] = ctx.get("model_cost", "0") return result ``` The variables available in the processing logic are: - `user_message` — the message typed by the tester, of type `str` - `history` — the full conversation history, of type `list[TypedDict[{"role": str, "content": str}]]` - `context` — the current session context, used to maintain state across turns - `collected_fields` — a `dict[str, str]` that can be mutated to track session-level statistics The return value must be a dictionary with the following keys: ```python { "output": "The text response shown to the tester", "context": ... # Updated context, can be a dict or JSON string } ``` ### Step 3 — Write the Instructions for User Impersonator The **Instructions for User Impersonator** field lets you define how an AI should simulate realistic user messages during a test session. Write a plain text prompt that describes the persona, context, and constraints the AI should follow when generating messages. ![Portal settings — Instructions for User Impersonator](portal-settings-user-impersonator.png) A good impersonator prompt describes: - The role the AI is playing (e.g., a banking customer, a patient) - The style and tone of messages it should generate - Any rules it should follow, such as staying in context or avoiding meta-commentary ### Step 4 — Configure the Advanced Configs The **Advanced Configs** section replaces the previous separate Test Definition and Closing Question schema editors. It accepts a single JSON object with two keys: `closingQuestions` and `testDefinition`. ![Portal settings — Advanced Configs](portal-settings-advanced-configs.png) Each field in either array follows this structure: ```json { "key": "field_key", "label": "Human readable label", "options": ["Option A", "Option B", "Option C"], "placeholder": "Placeholder text shown in the dropdown", "type": "selectbox" } ``` The `required` property is set directly on the field object and only needs to be included when set to `true`. Supported field types include `selectbox` for dropdown selection and `text` for free text input. Example for a classification pipeline: ```json { "closingQuestions": [ { "key": "is_predicted_intent_correct", "label": "Do you agree with the predicted intent?", "options": ["Yes", "No"], "required": true, "type": "selectbox" }, { "key": "correct_intent", "label": "If the above is No, what is the correct intent?", "options": [ "ACTIVATE CARD", "APPLY FOR LOAN", "BLOCK CARD", "CANCEL LOAN", "CANCEL TRANSFER", "CARD DETAILS" ], "type": "selectbox" }, { "key": "additional_comments", "label": "Any additional comments on why the classification was incorrect or could be improved?", "placeholder": "e.g. The message could also be interpreted as BLOCK CARD due to similar phrasing", "type": "text" } ], "testDefinition": [ { "key": "expected_intent", "label": "Expected Intent", "options": [ "ACTIVATE CARD", "APPLY FOR LOAN", "BLOCK CARD", "CANCEL LOAN", "CANCEL TRANSFER", "CARD DETAILS" ], "placeholder": "Select which intent this test is EXPECTED to classify", "type": "selectbox" } ] } ``` Supported field types for closing questions also include `qna`, which renders an interactive question-and-answer widget where testers can add multiple question-answer pairs. This is useful for conversational pipelines where testers want to record follow-up questions they would have asked. --- ## Using the Feedback Portal — Tester Guide ### Starting a Session Click **+ New Test** from the portal page to open a new session. The **Start a Conversation** screen will appear. ![Start a Conversation screen showing Test Definition, Pinned, Recently used, and Auto Chat tabs](session-start-conversation.png) If a Test Definition is configured, you can optionally select your expected values before typing your message. Use the **Pinned** and **Recently used** tabs to quickly reuse common test definitions. Then type your message in the chat input and press send to begin. ### During the Session The pipeline processes your message and returns its response in the chat view. ![Active chat session showing user message and pipeline response](session-chat-interface.png) You can rate individual responses using the thumbs up or thumbs down buttons. Continue the conversation if the pipeline is multi-turn, or move to closing if it is single-turn. When you are ready to finish, click the **End Session** link that appears below the last response. ### Ending the Session Clicking **End Session** opens the **Review Session** modal. ![Review Session modal — Feedback Summary, Testing Notes, and Closing Questions](session-review-modal.png) The modal shows: - **Feedback Summary** — a collapsible section summarising any thumbs down ratings from the session - **Testing Notes** — a free text field to document your overall observations - **Closing Questions** — a collapsible section with the structured questions defined in the Advanced Configs, with the pipeline's actual predictions shown inline in the question labels - **Mark all unmarked responses as 👍** — a checkbox to bulk-approve all responses that have not yet been rated Fill in your notes, answer the closing questions, and click **End Session** to submit. ### Auto Chat If the portal has **Instructions for User Impersonator** configured, you can use the **Auto Chat** tab to let an AI automatically generate and send user messages on your behalf. ![Auto Chat tab showing Max Turns and Instructions for User Impersonator](session-auto-chat-classification.png) Select the **Auto Chat** tab, set the **Max Turns**, and click **Start Auto Chat**. The AI will use the impersonator instructions as its prompt to generate realistic user messages and send them to the pipeline automatically. This is useful for quickly generating test sessions without manually typing each message. Once complete, a confirmation banner shows how many turns were run. The AI-generated messages appear in the chat view alongside the pipeline's responses, and you can proceed to end the session and fill in the closing questions as normal. ![Auto Chat completed — AI-generated message and pipeline response](session-auto-chat.png) --- ## Results and Insights ### Sessions Table All completed sessions are listed in the portal's results table. Each row represents one session and shows the session name, rating summary, status, testing notes, the date it was created, and any fields populated via `collected_fields`. ![Portal sessions list showing sessions from multiple contributors with status and testing notes](portal-sessions-list.png) Once a session is submitted, its status changes from **Ongoing** to **Completed** and the testing notes appear inline in the table. The **Created By** column shows which team member ran each session, allowing multiple contributors to provide feedback in parallel. Use the **Filter by status** dropdown to narrow the view to completed or ongoing sessions. ### Insights Click **Insights** in the top right of the portal to open the Insights view. It provides aggregate analysis across all sessions, organized into four tabs. #### Summary Shows high-level counts for the portal: total tests created, tests passed, tests failed, and sessions pending feedback. A Messages Summary section breaks down liked, disliked, and pending-feedback message counts. ![Insights — Summary showing Test Sessions Summary and Messages Summary](insights-summary.png) #### Contributors Shows how many tests each team member has run over time, with a bar chart of tests by date and a per-contributor breakdown. ![Insights — Contributors tab with tests over time chart](insights-contributors.png) #### Coverage Shows which Test Definition values have been exercised and which have not. A warning badge highlights options that have not yet been tested, helping you identify gaps in coverage. ![Insights — Coverage tab showing Expected Intent coverage chart](insights-coverage.png) #### Performance Shows overall test session performance over time, color-coded by outcome: all likes (green), containing dislikes (red), or no feedback (grey). Below the chart, a per-intent breakdown lets you drill into how performance varies across different expected values. ![Insights — Performance tab showing overall performance over time and per-intent breakdown](insights-performance.png) --- ## Tips for Portal Designers - Use `collected_fields` to surface any performance metrics your pipeline tracks — latency, cost, and token counts are especially valuable for benchmarking. - Write feedback instructions in plain text without markdown formatting for the best display in the portal UI. - Design closing questions to mirror your test definition fields so expected vs actual comparisons are easy to make in the results table. - Write the User Impersonator instructions to match the persona of your real end users — this ensures simulated messages are representative of actual traffic. --- ## Example — Customer Intent Classification Portal The following is a complete example of a Feedback Portal configured for a banking customer intent classification pipeline that classifies messages into 6 predefined intents: ACTIVATE CARD, APPLY FOR LOAN, BLOCK CARD, CANCEL LOAN, CANCEL TRANSFER, and CARD DETAILS. ### Instructions for User Impersonator ``` You are generating a message on behalf of a banking customer. Your task is to produce a single, realistic customer query that represents one of the following intents: ACTIVATE CARD, APPLY FOR LOAN, BLOCK CARD, CANCEL LOAN, CANCEL TRANSFER, or CARD DETAILS. The message must: 1. Be a natural, standalone customer query — not a continuation of a conversation 2. Clearly correspond to one specific intent without ambiguity 3. Be phrased the way an actual banking customer would write it 4. Not include any meta-commentary or explanations ``` ### Feedback Instructions ``` How to Use This Feedback Portal 1. Enter a user message in the chat input that represents a banking customer query. 2. Review the classified intent returned by the pipeline. 3. At the end of the session, answer the closing questions to confirm whether the predicted intent was correct. Tips for Good Test Cases - Use clear, single-intent messages to validate the pipeline's core classification accuracy - Test edge cases such as greetings, out-of-context messages, or human agent requests to check boundary behavior - Use realistic customer language — the way an actual banking customer would phrase their query - Try similar-sounding intents (e.g., BLOCK CARD vs CANCEL TRANSFER) to test the pipeline's ability to distinguish between closely related intents - Validate multilingual or informal phrasing to assess robustness Note: Each message is classified independently. The pipeline does not maintain conversation history across interactions. ``` ### Processing Logic ```python import json result = customer_intent_classification_pipeline(user_message, history=history, context=context) ctx = json.loads(result["context"]) collected_fields["intent_total_tokens"] = ctx.get("intent_total_tokens", "0") collected_fields["intent_latency_ms"] = ctx.get("intent_latency_ms", "0") collected_fields["intent_model_cost"] = ctx.get("intent_model_cost", "0") collected_fields["intent_input_cost"] = ctx.get("intent_input_cost", "0") collected_fields["intent_output_cost"] = ctx.get("intent_output_cost", "0") collected_fields["intent_tokens_per_second"] = ctx.get("intent_tokens_per_second", "0") return result ``` ### Auto Chat Since this pipeline classifies each message independently in a single turn, set **Max Turns** to **1** when using Auto Chat. This ensures the AI generates one realistic customer query per session, which the pipeline then classifies. ![Auto Chat configured with Max Turns set to 1 for single-turn classification](session-auto-chat-classification.png) ### Advanced Configs ```json { "closingQuestions": [ { "key": "is_predicted_intent_correct", "label": "Do you agree with the predicted intent?", "options": ["Yes", "No"], "required": true, "type": "selectbox" }, { "key": "correct_intent", "label": "If the above is No, what is the correct intent?", "options": [ "ACTIVATE CARD", "APPLY FOR LOAN", "BLOCK CARD", "CANCEL LOAN", "CANCEL TRANSFER", "CARD DETAILS" ], "type": "selectbox" }, { "key": "additional_comments", "label": "Any additional comments on why the classification was incorrect or could be improved?", "placeholder": "e.g. The message could also be interpreted as BLOCK CARD due to similar phrasing", "type": "text" } ], "testDefinition": [ { "key": "expected_intent", "label": "Expected Intent", "options": [ "ACTIVATE CARD", "APPLY FOR LOAN", "BLOCK CARD", "CANCEL LOAN", "CANCEL TRANSFER", "CARD DETAILS" ], "placeholder": "Select which intent this test is EXPECTED to classify", "type": "selectbox" } ] } ``` --- # Human Integrated Testing Source: https://docs.genguardx.ai/evaluate-and-approve/human-testing/ Markdown: https://docs.genguardx.ai/evaluate-and-approve/human-testing/index.md Description: Run human-integrated GGX tests where reviewers evaluate pipeline behavior, capture qualitative feedback, validate outcomes, and support approval decisions. The Human Integrated Testing module enables comprehensive testing of GenAI pipelines both before approval (Pre-Approval) and after deployment (Post-Approval/Post-Deployment), involving humans in the loop. ## What is a Pre-Approval Testing? The Pre-Approval Testing allows users to test the full end-to-end pipeline in a production-like environment, providing an opportunity to manually validate the final solution before deployment. It simulates real-world scenarios by replicating end-user experiences. This module abstracts the internal technical details, presenting only the final inputs and outputs for a streamlined evaluation process. **Note:** Currently, only the chat-based pipelines registered on the platform are available for manual testing. It enables reviewers to capture their feedback and scores during testing, offering valuable insights for the development team to improve the solution. Additionally, all testing data can be exported from the platform for different purposes, making it easier for developers and reviewers to collaborate. ## Performing Pre-Approval Testing on the Platform: All the chat-based pipelines are available for Pre-Approval Testing. 1. Click on **Start Session** in the Pre-Approval Testing module. 2. Choose the pipeline to be tested from a list of registered pipelines. 3. Once the pipeline is selected, a new window will open where you can start interacting. ![Testing Window](./test-session-output.png) 4. Start interacting in the chat window. Optionally, select a Persona to simulate real-life scenarios, or begin chatting directly. ![Persona Selection](./persona-selection.png) 5. Provide feedback: Use 👍 if satisfactory, 👎 if not; reasons/comments are prompted for 👎. 6. Once testing is completed, an experience summary can be recorded by providing an overall session rating and testing notes. ![Testing Feedback](./test-session-feedback.png) 7. You can download the transcript of the session for external documentation or further analysis. **Note:** More information about the Testing Session and GenAI Pipeline can be seen in the information panel on the left. There are 4 buttons in the information panel: - View other sessions from the same pipeline. - View the context being used in the pipeline. This is additional information that will be utilized by LLM to answer user questions. For example, it can be customer metadata being pulled from some knowledge base. - View more details about the pipeline and the Gen AI assets used in this pipeline. - View any configurations provided to the pipeline by the Modeler, such as temperature, random seed, etc. ![Information Panel](./test-session-information-panel.png) Every interaction, feedback, and comment in the Test Session is auto-saved immediately and recorded in history and recorded along with report cards in customized groups. ## Benefits of Pre-Approval Testing: - **Human-in-the-loop testing before approving** the pipeline for production. - Easy to follow and **intuitive user interface** to test pipelines. - **Abstraction** of complex pipeline logic from MRM/Fair-Lending/Business teams. - **Record feedback and generate report cards** for continuous improvement. - Allows **Multi-Turn testing** for real-world scenarios. - **Load transcripts from other environments** for human scoring. Similar to this a pipeline can be tested/monitored post approval also. Visit [Deploy and Monitor](../../deploy-and-monitor/) section to know more on this. --- # Reporting Source: https://docs.genguardx.ai/evaluate-and-approve/reporting/ Markdown: https://docs.genguardx.ai/evaluate-and-approve/reporting/index.md Description: How reports are built in GGX — metadata, data schema, parameter schema, computation, and visualization — plus how to register, test, and reuse them. ## What is a Report? A report is an analytical entity that turns the output of a job into insight. It reads a data source, applies computation logic (statistics, grouping, filtering, model-based scoring), and presents the result through visual elements such as charts, tables, and narrative summaries — helping stakeholders make data-driven decisions. Once registered, a report becomes a **reusable asset**: the same report can be run against many objects (Pipelines, Models, RAGs, Prompts) and combined with others into use-case-specific **dashboards**. ## Anatomy of a report Every report is made of **five building blocks** — three that *define* what the report expects, and two that hold the *logic* it runs.
![The five building blocks of a report: Metadata, Data Schema, and Parameter Schema define the report; Computation and Visualization hold its logic.](./report-anatomy.svg)
Metadata, Data Schema, and Parameter Schema describe the report; Computation and Visualization do the work.
### 1. Metadata Metadata is the descriptive header of the report — how it is named, classified, and discovered in the Report Registry. Only the **Name** is required up front; the classification fields can be left empty and filled in later. | Field | Required | What it captures | |-------|----------|------------------| | **Name** | | Human-readable name of the report. | | **Object Type** | | The kind of object the report runs on (Foundation Model, Pipeline, RAG, Prompt, …). | | **Description** | optional | What the report measures and how to read it. | | **Risk Type** | optional | The risk category the report addresses (e.g. Accuracy, Bias, Toxicity, Vulnerability). | | **Task Type** | optional | The task being evaluated (e.g. classification, summarization, retrieval, generation). | | **Risk Domain** | optional | The governance lens — MRM, Fair Lending, Business/Tech, and so on. | | **Evaluation Methodology** | optional | How the metric is computed: **LLM-as-a-Judge**, **NLP / ML Algorithms**, **Rule-based**, or **Others**. | ### 2. Data Schema The Data Schema declares the **columns the report needs** from the job output. It is the contract between a report and the objects it can run on: any object whose output provides these columns can use the report. Each entry has a column name, a mandatory flag, and a description. ```ts { columnName: string; // the column the report reads mandatory: boolean; // must the data provide it? description: string; // what the column holds } ``` For example, a response-accuracy report might expect: | Column | Mandatory | Description | |--------|-----------|-------------| | `output` | Yes | The response produced by the object under test. | | `expected_answer` | Yes | The ground-truth answer to compare against. | | `context` | No | Retrieved context passed to the model (used for grounding checks). | :::tip[Design for reuse] Write the schema against **generic** column names (*output*, *context*, *expected_answer*) rather than object-specific ones. A report expressed in generic columns can run across many agents and tables; a report hard-wired to one object's columns cannot. ::: ### 3. Parameter Schema The Parameter Schema lists the **inputs a user supplies at run time** — thresholds, category choices, the evaluator model to use, and so on. The platform renders each parameter as a form control based on its `type`, and the values arrive in the computation as `job.parameters`. ```ts { key: string; // name used in code, e.g. job.parameters["threshold"] label: string; // shown in the run form type: 'numberbox' | 'textbox' | 'selectbox' | 'selectbutton'; placeholder: string; description: string; defaultValue: string; options: string[]; // choices for selectbox / selectbutton required: boolean; }[] ``` | `type` | Renders as | Use for | |--------|-----------|---------| | `numberbox` | Numeric input | Thresholds, top-k, temperatures (e.g. `0.5`). | | `textbox` | Free-text input | Labels, column names, free-form values. | | `selectbox` | Dropdown | One choice from a longer list (`options`). | | `selectbutton` | Button group | One choice from a short set of `options`. | For example, a single pass/fail threshold: ```json [ { "key": "threshold", "label": "Pass threshold", "type": "numberbox", "placeholder": "0.7", "description": "Minimum score for a response to count as correct.", "defaultValue": "0.7", "options": [], "required": true } ] ``` :::caution[Handle optional parameters] If a parameter is **not** required, the computation may receive it as empty or missing — read it defensively (fall back to `defaultValue`) so the report still runs. ::: ### 4. Computation The computation is where the report does its work. The platform hands it two values: - **`data`** — the job output as a **DuckDB DataFrame**. It always holds the object's **input columns** plus its **output**. What the output looks like depends on the object: a chat/LLM object, for instance, emits an `output` string and any `context` the author chose to surface (`{"output": str, "context": ...}`). - **`job`** — a handle to the run and its context: `job.parameters` (the values from the Parameter Schema), `job.current` (the object under test), and `job.current.name` (its name). You can operate on `data` with DuckDB directly, or call `data.df()` to convert it to a **pandas DataFrame** and work in pandas — the logic below is written the same way either way. Whatever the computation produces is passed to the visualization step as **`raw_output`**. It can be a single DataFrame or a **dict of DataFrames**: ```python # `data` (a DuckDB DataFrame) and `job` are provided by the platform. df = data.df() # convert to pandas threshold = float(job.parameters["threshold"] or 0.7) # Score each row, then summarize. scored = df.assign( correct=lambda d: similarity(d["output"], d["expected_answer"]) >= threshold ) summary = ( scored.groupby("correct").size() .rename("count").reset_index() ) # The value of the computation is exposed to visualization as `raw_output`. raw_output = {"scores": scored, "summary": summary} ``` ### 5. Visualization Each visualization runs on the output of the computation, available as **`raw_output`**. A report can have several outputs — add more with the **Add Output** button — and each returns one of: - a **Plotly figure** — charts and interactive graphics; - a **pandas / DataFrame** — rendered as a sortable, filterable **grid table** with no extra code; - an **HTML / Markdown** string — narrative summaries, executive overviews, formatted callouts. ```python import plotly.express as px # `raw_output` and `job` are provided by the platform. summary = raw_output["summary"] fig = px.bar( summary, x="correct", y="count", title=f"{job.current.name} — response accuracy", ) fig # returning a Plotly figure renders it as a chart ``` A metric-style output simply returns a numeric value.
![At run time the job output and the user's parameters feed the computation, which produces raw_output; the visualization turns raw_output into figures that assemble into a dashboard.](./report-dataflow.svg)
At run time: job data and parameters feed the computation; its raw_output feeds each visualization; the figures assemble into a dashboard.
## Benefits of report registration - **Customize reports** to meet MRM, Fair Lending, Business, and Development requirements. - Combine multiple reports into **use-case-specific dashboards** while running jobs. - **Track all modifications** with enhanced version management. - **Collaborate and approve** with MRM, FL, and other team members. - Track usage through the **Lineage Tracking** capability. - **Reuse across cases** to avoid repeating the approval process. - Make **quick updates** as needs and business logic change. ## Registering a report The **Report Registry** organizes every report in one place for easy tracking, monitoring, and creation. To register a new one, go to **Resources → Reports** and click **Create**, then work through the five building blocks: 1. **Fill in the metadata.** Give the report a **Name** and choose the **Object Type** it runs on. Add a **Description**, and — now or later — the **Risk Type**, **Task Type**, **Risk Domain**, and **Evaluation Methodology**. Optionally assign a **Group** (for organization) and an **Approval Workflow** (the chain for moving from Draft to Approved). 2. **Declare the Data Schema.** List the columns the report reads, marking each mandatory or optional. These are the columns the object's output must provide. 3. **Declare the Parameter Schema.** Add the run-time inputs, choosing a control type (`numberbox`, `textbox`, `selectbox`, `selectbutton`), a default, options where relevant, and whether each is required. 4. **Provide example data.** Attach a sample dataset — from a registered table or an uploaded file — so the report can be run and validated during authoring. 5. **Write the computation.** Read `data` and `job`, compute your metrics and tables, and return them (a DataFrame or a dict of DataFrames). Select any additional resources (Global Functions, Models, evaluator LLMs) the logic needs. 6. **Write the visualization(s).** Turn `raw_output` into figures — Plotly charts, grid tables, or HTML/Markdown. Use **Add Output** for more than one figure. 7. **Add notes or attachments**, then click **Create** to finalize registration. :::tip[Plan the dashboard first] Sketch the dashboard before you build. Decide which figures you need and how they sit together — a common layout is an HTML **executive summary** on top, per-metric **deep-dive** charts in the middle, and a **raw assessment** grid table at the bottom for export. ::: ## Testing a report with a Simulation job A report can be validated on its own before it is trusted for governance. Run it through a **[Simulation](../simulation/)** job with representative data and known expected outputs, then confirm the numbers and visuals come out as intended. Because the report is just an asset, the same run doubles as evidence you can attach to an [Approval Workflow](../approval-workflows/). Once validated, the report is ready to be selected when running jobs on other objects — the dashboard is produced automatically when the job completes.
## Reusing reports across agents Because a report is written against the columns it expects (for example *user message*, *response*, and *context*), a single generic report can run across many agents or tables, while more specific reports capture bespoke metrics for one agent pattern — single-turn, multi-turn, RAG, and so on. A common layout is an HTML **executive summary** at the top, per-metric **deep-dive** figures, and a **raw assessment** table that stakeholders can export for further analysis. :::note[Out-of-the-box reports] You do not have to start from a blank page. GGX ships with a library of ready-made reports and metrics — covering common needs such as response accuracy, intent accuracy, stability, and vulnerability — that you can run as-is or copy and adapt. Domain starter kits bundle the reports and metrics most relevant to sectors like financial services and healthcare. ::: --- # Simulation Source: https://docs.genguardx.ai/evaluate-and-approve/simulation/ Markdown: https://docs.genguardx.ai/evaluate-and-approve/simulation/index.md Description: Run GGX simulation jobs on datasets to evaluate objects, generate reports, inspect metrics, validate evaluators, and compare results before approval. The most common type of execution is a **Simulation** - used to execute analytics contained in the definition of an object that has been registered in the platform. And additionally, run dashboards and reports on that output. A Simulation task typically involves: - **Current object:** Object which is currently selected and being tested. This can be a Pipeline, Model, RAG, or Prompt. - **Data Source:** The data to run the object on. - **Report & Metrices:** Exact evaluation metrics to be run on the output. ## How to run a Simulation Task? - **Register the object** on the platform. - Go to the **Details** page of an object to be compared and click on the **Run -> Simulation** button. - **Provide description** about the run. - **Select Dashboard** to be evaluated in **Dashboard Selection**. - Prepare the data in **Data Sources** which will be used for evaluation. - Click on **Run** at the bottom and wait for job completion. - Once a job has been submitted it starts in the NEW status - The job will go through the following statuses: COMPILING > QUEUED > RUNNING and finally stop at COMPLETED of FAILED All simulation tasks are systematically recorded on the platform and displayed on the Jobs page of an object in a structured format. They can also be exported as part of the automated documentation process. > **Note:** The platform allows customizing reports and dashboards specifically for simulation tasks. > **Note:** GGX allows running jobs in parallel and multiple threads at a time within a job to expedite the evaluation process. ## Validating an evaluator Simulation is also how you check that an *evaluator* can be trusted. An LLM-as-a-judge is itself a registered model, so you can run it over a **ground-truth dataset** and compare its verdicts against the known answers — computing classification metrics just as you would for any model. The same applies to a whole report: run it on labelled data with expected outputs and confirm the numbers and visuals come out as intended before relying on it in monitoring. These runs can then be attached to the object's [Approval Workflow](../approval-workflows/) as validation evidence. --- # Frequently Asked Questions Source: https://docs.genguardx.ai/faq/ Markdown: https://docs.genguardx.ai/faq/index.md Description: Common questions about evaluating and monitoring GenAI applications in GGX — production monitoring, reports and dashboards, LLM-as-a-judge metrics, reusable assets, data and integrations, governance and testing, versioning, and human feedback. This page answers the questions teams ask most often when they start evaluating and monitoring GenAI applications on GGX. Each answer links to the reference page that covers it in depth. ## Production monitoring ### How do I monitor a GenAI agent running in production? Three building blocks, in order: 1. Bring your production interaction data into a registered **Data Table** (upload a file, or connect to a database, cloud bucket, or data lake). 2. Build one or more **Reports** that compute the evaluation metrics you care about. 3. Create a **monitoring Simulation** that runs those reports over the table and publishes a **Dashboard**. The dashboard is the end result — an executive summary plus deep-dive sections you can drill into. ### What is the difference between pre-production and post-production evaluation? **Pre-production** tests a registered object (a Pipeline, Model, RAG, or Prompt) on curated test data *before* approval. **Post-production** monitors *live* data after deployment. The same reports and metrics can serve both — a report's **Object Types** let you mark it eligible for pipelines (pre-prod) and for monitoring (post-prod), so you write the logic once. ### How do I keep a monitoring view continuously up to date? Set the job up as a **recurring simulation** and select its first iteration for the **Metrics Dashboard**. The dashboard then tracks the latest completed run, and the platform can notify you as iterations near completion. ### Can I filter or reshape data before it reaches a report? Yes. The monitoring job has a pre-processing step between the raw table and the report where you can filter rows, rename columns, sample, or group records — reusing registered **Global Functions** so you do not rewrite the logic per agent. Filtering can also live inside the report's computation logic. ## Reports and dashboards ### What goes into a report? A report is made of five building blocks: **metadata** (name, object type, risk/task classification, evaluation methodology), a **data schema** (the columns it reads), a **parameter schema** (run-time inputs), **computation logic** (filtering, transformations, and metric calculations), and **visualization logic**. The job's data arrives as a dictionary of DuckDB DataFrames that you can compute over in Pandas. ### What kinds of visualizations can a report produce? Markdown/HTML for narrative summaries, and **Plotly, Seaborn, or Matplotlib** for charts. A returned Pandas DataFrame renders automatically as a sortable, filterable grid table. A common pattern is an HTML **executive summary** at the top, followed by per-metric deep-dive figures and a raw assessment table you can export. ### Can one report be reused across many agents? Yes — that is the point of registering it. Write a report generically against the columns it expects (for example, *user message*, *response*, and *context*) and it runs on any table or agent that provides them. You can also build agent-specific reports when a use case needs bespoke metrics. ### Do I need to define metrics separately for every agent? No — you generalise by **agent pattern**, not per agent. Conversational (single-turn), multi-turn, and RAG agents each have their own natural set of metrics, so you build one report per pattern and run it across every agent of that type, adding agent-specific metrics only where a use case truly needs them. ### Does GGX come with ready-made reports and metrics? Yes. A library of out-of-the-box reports and metrics ships with the platform — covering common needs such as response accuracy, intent accuracy, stability, and vulnerability — and domain starter kits bundle reports relevant to sectors like financial services and healthcare. You can run them as-is or copy and adapt them. ### Can different stakeholders each see their own dashboard? Yes. Dashboards and custom views can be tailored to an audience and surfaced by **role** — an environment-wide executive summary, a team or product-domain view, a per-customer/tenant view, or an object-level view for the developers of a single agent. ### How do I build a full dashboard? Group several reports together — each contributing its own metrics and visuals — into a use-case dashboard that is published when the job completes. ## Evaluation metrics and LLM-as-a-judge ### What is an LLM-as-a-judge? It is an LLM registered as a **Model** in the Model Catalog whose job is to *score* another model's output against criteria defined in a **Prompt** — for example, rating answer relevancy on a 0–4 scale. Because it is a registered asset, the same judge is reusable as a resource across any report. ### Which evaluation metrics can I capture? Common ones include **answer relevancy**, **factual accuracy / faithfulness**, **groundedness**, **context relevancy**, **completeness**, **toxicity**, **PII detection**, and **tool-selection / tool-call accuracy** — alongside operational metrics like **latency, cost, and token counts**. You can mix LLM judges, rule-based checks, and NLP libraries to compute them. ### How do I know I can trust a judge? Test the judge the same way you test any model: run it over a **ground-truth dataset** and compute classification metrics to see how well it agrees with known answers, then refine its prompt (including via Prompt Optimization) to fit your use case before relying on it. ### When should I use a heuristic check instead of an LLM judge? Use cheap, deterministic **heuristics** first — thresholds on tokens, length, or cost; staleness checks; keyword-based failure detection; PII detection via libraries. They are instant and free, and they triage the obvious cases. Reserve LLM judges for the nuanced quality questions that rules cannot answer. ### How is PII detected — does it always need an LLM? No. PII is usually caught with **deterministic rule-based / NLP libraries** rather than an LLM — fast and inexpensive — and the same tools can redact the detected values. LLM-based checks are reserved for nuanced, context-dependent cases, such as distinguishing data the user explicitly asked for from a genuine leak. ### Can I replay production prompts to choose a better model? Yes — that is a **Comparison** job. Register challenger versions (swap in a different model, prompt, or component), run them against a common dataset, and compare the metrics side by side to select the best candidate. ## Reusable assets ### How does GGX avoid rewriting the same evaluation logic everywhere? Everything is a reusable, governed asset: **Models** (including judges), **Prompts** (templates), and **Global Functions** (utilities). You compose these resources into a report rather than copy-pasting code, and registration brings change tracking, approvals, and lineage with it. ### How do I add a utility like cosine similarity, deduplication, or PII redaction? Register it as a **Global Function**. It is then callable both inside a report's computation logic and in the monitoring pre-processing step — write it once, reuse it everywhere. ### Can I author assets in my own IDE? Yes. Use the GGX Sync package to develop in your editor and push assets into the platform, keeping source control and review in your normal workflow. ## Data and integrations ### My production export is huge — do I have to upload a CSV? No. For large datasets, connect GGX to a **database, cloud bucket, or data lake** (for example, a Parquet source) instead of uploading through the UI. The data is then fetched server-side and still viewable as a table. ### How do I connect an LLM provider or API key? Go to **Settings → Platform Integrations**, pick the provider, add its credentials, and use **Test Connection** to confirm. The **Advanced** tab lets you define environment variables available across the platform, and you can request models that are not yet in your inventory. ### Can I swap an embedding or judge model later? Yes — assets are plug-and-play. You can start with an open-source embedding model and switch to a hosted provider with no change to the surrounding logic, because the function simply expects text in and an embedding (or score) out. To choose between candidates, run them over a labelled dataset and compare. ### What data can I run a report or simulation on? Several sources: a **registered table** (uploaded, or connected from a data lake/bucket), a **custom file** uploaded for the run, **AI-generated** data when you need to synthesize cases you do not have, or **human-labelled** data promoted from an Annotation Queue for ground-truth runs. ## Governance, testing, and versioning ### Do I still need Git to manage this evaluation code? For governance, the platform replaces what teams usually bolt onto Git: every change to a report, judge, or function is snapshotted with author and reason, and promotion runs through an approval workflow that non-developers can read. You can still develop in your own editor and sync the code in — but you do not need Git to get change history, approvals, or an audit trail. ### How do I test changes before they are approved? Several ways: use **Test Code** while editing to run logic against sample inputs; run a **Simulation** or **Comparison** over test or ground-truth data; or export the code and test it in your own environment. As part of approval, a **regression suite** can be re-run automatically and external **CI** checks can be required before promotion. ### Can I see which version of a report produced a past dashboard? Yes. Every object keeps a full **Change History** of snapshots, so from a dashboard you can open the exact version of the report that generated it at that point in time — and restore or compare versions as needed. ## Human feedback and continuous improvement ### How do subject-matter experts give feedback on agent quality? Through **Feedback Portals** — an SME interacts with the pipeline directly, rates individual responses, and answers structured closing questions. The portal aggregates this into insights on coverage and performance. ### How do I label production data or build ground truth? Use **Annotation Queues** to label live production outcomes (Is Accurate, Notes, Ground Truth) in an auditable, deduplicated workflow — a major improvement over spreadsheets. ### How do I turn recurring feedback into new evaluations? This is the loop that makes evaluation scale: collect SME feedback and labels, identify the recurring failure categories, codify each category as a new **judge / report**, then run it across *all* your agents on a recurring schedule. New problems become new automated checks over time. ### How do I get alerted when something goes wrong? Configure **Automated Alerts** on your dashboard data — conditions and thresholds with a severity level. A triggered alert can send notifications, **email or Slack** messages, raise a flag, or open a review on the responsible object. This is how you catch runaway behaviour early — for example, an agent burning an abnormal number of tokens. --- # GGX Integrations Source: https://docs.genguardx.ai/integrations/ Markdown: https://docs.genguardx.ai/integrations/index.md Description: Connect GGX with LLM providers, LLM gateways, observability tools, agent frameworks, report providers, voice providers, and SSO systems across the AI lifecycle. GGX is built with extensibility in mind and supports a wide range of integrations that allow you to seamlessly plug into existing workflows and tools. These integrations are organized into the following categories: 1. **LLM Providers:** Connect to popular foundational models from leading providers to power your generative experiences. Examples: [OpenAI](https://openai.com/), [Google Vertex AI](https://cloud.google.com/vertex-ai),[Azure AI](https://azure.microsoft.com/),[Amazon Bedrock](https://aws.amazon.com/bedrock/), [DeepSeek](https://www.deepseek.com/), [Anthropic](https://www.anthropic.com/), [HuggingFace](https://huggingface.co/), [Nvidia NIM](https://www.nvidia.com/en-us/ai/), [GitHub Models](https://github.com/marketplace/models) 2. **LLM Gateways:** Register an LLM gateway in the GGX Model Registry, then route GGX prompts, RAGs, pipelines, simulations, and monitoring jobs to any LLM the gateway has access to. Examples: [LiteLLM](llm-gateways/litellm/), [Portkey](llm-gateways/portkey/), [OpenRouter](llm-gateways/openrouter/), [Cloudflare AI Gateway](llm-gateways/cloudflare-ai-gateway/), [Databricks AI Gateway](llm-gateways/databricks-ai-gateway/) 3. **Observability:** Connect AI traces, logs, scores, alerts, and review signals to GGX so monitoring feeds judges, human review, ground truth, bug categorization, and lifecycle governance. Examples: [LangSmith](observability/langsmith/), [Arize Phoenix](observability/arize-phoenix/), [Langfuse](observability/langfuse/), [Humanloop](observability/humanloop/), [Datadog](observability/datadog/) 4. **Agent Providers & Frameworks:** Leverage pre-built agent providers or bring your own orchestration frameworks to create and manage intelligent, multi-step agent workflows with minimal setup. Examples: [Vertex AI Agent Playbooks](https://cloud.google.com/dialogflow/cx/docs/concept/playbook), [AgentForce (Salesforce)](https://www.salesforce.com/in/agentforce/), [Microsoft Copilot Studio](https://www.microsoft.com/en-us/microsoft-365-copilot/microsoft-copilot-studio), [Vapi AI](https://vapi.ai/),[GitHub Copilot](https://github.com/features/copilot), [Amazon Lex](https://aws.amazon.com/lex/), [CustomGPT](https://customgpt.ai/) 5. **Report Providers:** Plug in evaluation tools to assess your data, RAGs, pipelines, agents, etc. effectively. Examples: [CleanLabs](https://cleanlabs.ai/), [Perspective API](https://perspectiveapi.com/) 6. **Voice Providers:** Setup integrations to services that provide speech-to-text, text-to-speech, Voice Agents Examples: [Deepgram](https://deepgram.com/) 7. **Single Sign-On (SSO) Integrations:** Setup integrations to services for seamless Authentication and Authorization Examples: [Google Workspace](https://workspace.google.com/), [Auth0 (by Okta)](https://auth0.com/), [WorkOS](https://workos.com/), [OneLogin](https://www.onelogin.com/) ## Observability and closed-loop monitoring AI observability platforms are valuable because they collect production traces, scores, feedback, and alerts. GGX makes those signals operational across the AI lifecycle. Traces from tools such as LangSmith, Arize Phoenix, Langfuse, Humanloop, and Datadog can be connected to GGX. Once connected, GGX can run automated judges and route selected interactions to human review. Positive reviews can be promoted into ground truth. Negative reviews can become bugs or findings. Repeated issues can be grouped into common failure categories. This reduces the review funnel and makes monitoring useful beyond dashboards. Monitoring evidence starts to power better datasets, better judges, clearer approval evidence, and more targeted fixes before the next deployment. --- # Cleanlab Source: https://docs.genguardx.ai/integrations/evaluation-providers/cleanlab/ Markdown: https://docs.genguardx.ai/integrations/evaluation-providers/cleanlab/index.md Description: Connect Cleanlab with GGX to use external data quality and evaluation signals in model, RAG, pipeline, and dataset assessment workflows. ![alt text](cleanlab_logo.png) ## About Cleanlab Pioneered at MIT and proven at 50+ Fortune 500 companies, Cleanlab provides software to detect and remediate inaccurate responses from Enterprise AI applications. Cleanlab detection serves as a trust and reliability layer ensuring Agents, RAG, and Chatbots remain safe and helpful. Recognized among the **Forbes AI 50, CB Insights GenAI 50, and Analytics India Magazine’s Top AI Hallucination Detection Tools**, the company was founded by three MIT computer science PhDs and is backed by $30M investment from Databricks, Menlo, Bain, TQ, and the founders of GitHub, Okta, and Yahoo. ## Integrating Cleanlab Simply enter your Cleanlab API key once in the **Integrations** section of GGX. This enables authorized users to access Cleanlab’s capabilities within the platform. Once integrated, Cleanlab can be used as any other python package on the platform. ``` from cleanlab_tlm import TLM tlm = TLM() trustworthiness_score = tlm.get_trustworthiness_score("", response="") ``` ## Potential Usage Within GGX Cleanlab strengthens GGX by adding a **trust and reliability layer** that evaluates the accuracy, relevance, and safety of agents and their underlying components. With Cleanlab, one can systematically identify hallucinations, off-topic responses, unsafe outputs, and other critical issues across agents, RAG pipelines, and broader LLM workflows. Cleanlab’s metrics and scores can be used to create insightful reports for validating systems and can also serve as guardrails to ensure your agents respond safely. ### Key Use Cases - **RAG System Evaluation** Automatically detect inaccuracies, assess retrieval quality, and surface knowledge gaps in RAG pipelines. - **Agent & Pipeline Response Evaluation** Cleanlab’s TLM (Trustworthiness Language Model) assigns confidence scores to LLM responses, flagging hallucinations, ambiguous answers, and unsafe content with detailed explanations. - **Data Quality & Reliability** Identify and resolve issues such as mislabeled data, ambiguous examples, or statistical outliers ensuring your models are trained or validated on clean, trustworthy data. ### Example Use Case: Enhancing IVR System Evaluation with Cleanlab Cleanlab significantly helps test IVR (Interactive Voice Response) systems by providing trustworthiness scores for various AI-driven processes. It evaluates the reliability of LLM decisions in classifying caller intent and routing calls, providing responses. Furthermore, Cleanlab can assesses the accuracy of human data labels used in validation, helping to pinpoint mislabeled data. For the auto-generated call summaries and candidate responses, Cleanlab's scores can highlight problematic outputs and areas where the agent may need further improvement. ![alt text](raw_trust_score.png) For example, check below, Cleanlab can score how trustworthy each response is, making it easy to spot which types of questions or categories have issues. ![alt text](trust_score.png) ## Want to Learn More ? - [Cleanlab Codex Documentation](https://help.cleanlab.ai/codex/) - [Cleanlab TLM Overview](https://help.cleanlab.ai/tlm/) - [Cleanlab Studio Overview](https://help.cleanlab.ai/studio/quickstart/) --- # LLM Gateways Source: https://docs.genguardx.ai/integrations/llm-gateways/ Markdown: https://docs.genguardx.ai/integrations/llm-gateways/index.md Description: Use LLM gateways with GGX by registering gateway-backed models in the Model Registry and routing requests to any LLM the gateway can access. LLM gateways provide a common API layer in front of one or more model providers. They are useful when your organization already centralizes routing, provider credentials, cost controls, logging, rate limits, caching, fallback, or provider selection outside GGX. In GGX, an LLM gateway can be registered in the [Model Registry](../../register-and-refine/inventory-management/model-catalog/) as a Model. Once registered, that model can be used in prompts, RAGs, pipelines, simulations, comparisons, approval workflows, and production monitoring just like any directly integrated LLM provider. ## How the pattern works 1. Configure the LLM gateway with the providers and model aliases it is allowed to access. 2. Store the gateway base URL, API key, and default model alias in GGX environment variables or integration settings. 3. Register a GGX Model that calls the gateway endpoint. 4. Use the registered GGX Model in pipelines and evaluations. 5. Let the gateway route to any underlying LLM it has access to, while GGX records governance metadata, test evidence, approvals, lineage, and monitoring results. ## Registration options | Option | When to use it | | --- | --- | | **One GGX Model per gateway model alias** | Use this when reviewers should approve and monitor each underlying LLM separately. For example, register `gateway_gpt4o`, `gateway_claude_sonnet`, and `gateway_gemini_flash` as separate GGX Models. | | **One parameterized GGX Model** | Use this when the gateway chooses the model dynamically or when the caller should pass a `model` argument. This is useful for routing tests, fallback tests, and cost or latency comparisons. | | **Gateway-backed pipeline** | Use this when the gateway is only one component in a larger GGX Pipeline that also includes prompts, retrieval, guardrails, parsing, or business logic. | ## Model Registry setup When registering an LLM gateway in the Model Registry, capture enough metadata for business, risk, and audit reviewers to understand what is behind the gateway. | Field | Recommended value | | --- | --- | | **Name** | A clear name such as `litellm_gateway_gpt4o`, `portkey_gateway_default`, or `cloudflare_gateway_openai`. | | **Description** | State that the GGX Model calls an LLM gateway and identify the gateway, allowed providers, model aliases, and routing policy. | | **Input Type** | Use an API-based or Python-function model implementation, depending on how the gateway is reached in your environment. | | **Arguments** | Include `messages`, `prompt`, `model`, `temperature`, `max_tokens`, and any request tags your gateway supports. | | **Output Type** | Use `String` for simple responses or `Map[String, String]` when returning response text plus metadata such as selected model, gateway request ID, or routing outcome. | | **Risk Assessment** | Document the gateway owner, approved provider list, fallback policy, logging policy, data retention behavior, and whether prompts or responses are stored by the gateway. | ## Generic OpenAI-compatible example Many gateways expose an OpenAI-compatible API. Register a GGX Model with scoring logic like this, then set the environment variables for the gateway you use. ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("LLM_GATEWAY_API_KEY"), base_url=os.getenv("LLM_GATEWAY_BASE_URL"), ) selected_model = model if model else os.getenv("LLM_GATEWAY_MODEL") completion = client.chat.completions.create( model=selected_model, messages=messages, temperature=float(temperature), max_tokens=int(max_tokens), ) return { "output": completion.choices[0].message.content, "model": selected_model, } ``` ## Gateway-specific guides - [LiteLLM](litellm/) - [Portkey](portkey/) - [OpenRouter](openrouter/) - [Cloudflare AI Gateway](cloudflare-ai-gateway/) - [Databricks AI Gateway](databricks-ai-gateway/) ## What GGX adds on top of an LLM gateway LLM gateways centralize runtime access. GGX adds lifecycle governance around the AI system that uses that gateway: - Model Registry inventory, ownership, and descriptions. - Version history and lineage for prompts, models, RAGs, pipelines, reports, and data. - Automated simulations and comparisons across gateway-backed models. - Standardized risk reports for accuracy, stability, bias, toxicity, vulnerability, hallucination, retrieval quality, and other compliance dimensions. - Approval workflows for business, risk, legal, compliance, technology, and model-risk reviewers. - Production monitoring, dashboards, alerts, and annotation queues. --- # Cloudflare AI Gateway Source: https://docs.genguardx.ai/integrations/llm-gateways/cloudflare-ai-gateway/ Markdown: https://docs.genguardx.ai/integrations/llm-gateways/cloudflare-ai-gateway/index.md Description: Register Cloudflare AI Gateway as a GGX Model Registry model and route GGX requests to providers available through your Cloudflare gateway. [Cloudflare AI Gateway](https://developers.cloudflare.com/ai-gateway/) provides visibility and control for AI applications, including analytics, logging, caching, rate limiting, retries, and model fallback. Cloudflare supports provider-specific endpoints that preserve the provider API schema while adding AI Gateway features. In GGX, Cloudflare AI Gateway can be registered in the [Model Registry](../../register-and-refine/inventory-management/model-catalog/) as a Model. The registered GGX Model calls the Cloudflare gateway endpoint, and Cloudflare can connect to any provider or model that your gateway configuration and account permissions allow. ## When to use this integration Use Cloudflare AI Gateway with GGX when: - Cloudflare is the control plane for AI traffic observability, caching, rate limiting, retries, or fallback. - Your production application already routes provider traffic through Cloudflare AI Gateway. - You want GGX evaluations to test the same gateway path used in production. - You want Cloudflare to handle runtime controls while GGX handles inventory, simulations, approvals, compliance evidence, and monitoring workflows. ## Register Cloudflare AI Gateway in the Model Registry | GGX setting | Recommended value | | --- | --- | | **Name** | `cloudflare_gateway_` | | **Description** | Identify the Cloudflare account, gateway ID, provider endpoint, model alias, and runtime controls such as caching, rate limits, retries, or fallback. | | **Model Provider** | Use a custom/API-based model or Python function that calls the Cloudflare AI Gateway endpoint. | | **Arguments** | `messages`, `model`, `temperature`, `max_tokens` | | **Environment variables** | `CLOUDFLARE_AI_GATEWAY_BASE_URL`, `CLOUDFLARE_AI_GATEWAY_API_KEY`, `CLOUDFLARE_AI_GATEWAY_MODEL` | Cloudflare provider-specific endpoints use this pattern: ```txt https://gateway.ai.cloudflare.com/v1/{account_id}/{gateway_id}/{provider} ``` For OpenAI-compatible providers, configure `CLOUDFLARE_AI_GATEWAY_BASE_URL` to the provider-specific Cloudflare endpoint for your account and gateway. ## Example scoring logic ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("CLOUDFLARE_AI_GATEWAY_API_KEY"), base_url=os.getenv("CLOUDFLARE_AI_GATEWAY_BASE_URL"), ) selected_model = model if model else os.getenv("CLOUDFLARE_AI_GATEWAY_MODEL") completion = client.chat.completions.create( model=selected_model, messages=messages, temperature=float(temperature), max_tokens=int(max_tokens), ) return { "output": completion.choices[0].message.content, "model": selected_model, } ``` ## Governance notes - Register separate GGX Models when Cloudflare routes to different providers or when each provider requires separate review. - Document Cloudflare caching, retry, rate-limit, and fallback settings in the GGX Model Risk Assessment. - If Cloudflare fallback is enabled, include route metadata in GGX output or monitoring data when available. - Use GGX monitoring to evaluate production behavior after Cloudflare has applied runtime controls. References: - [Cloudflare AI Gateway overview](https://developers.cloudflare.com/ai-gateway/) - [Cloudflare AI Gateway get started](https://developers.cloudflare.com/ai-gateway/get-started/) --- # Databricks AI Gateway Source: https://docs.genguardx.ai/integrations/llm-gateways/databricks-ai-gateway/ Markdown: https://docs.genguardx.ai/integrations/llm-gateways/databricks-ai-gateway/index.md Description: Compare Databricks Unity AI Gateway with GGX for GenAI governance, automated compliance testing, approval evidence, production monitoring, and risk management. Databricks Unity AI Gateway and GGX solve adjacent governance problems. Databricks focuses on governing AI service access and runtime traffic inside the Databricks ecosystem. GGX focuses on the end-to-end Responsible AI lifecycle: registering GenAI assets, testing risk and business performance, producing approval evidence, monitoring production behavior, and feeding findings back into refinement. ## Executive summary Databricks AI Gateway is the **runtime AI service control plane** for Databricks-centered deployments. It answers: - Who can call this AI service? - Which model or MCP service does traffic reach? - How much usage or cost is allowed? - Which request or response policies should be enforced? - What happened at the request, token, latency, cost, and payload level? GGX provides the **Responsible AI lifecycle management**. It answers: - What stage of the lifecycle is the GenAI use case, agent, etc. at - Planning, Development, Deployed? - Which risks matter for this use case? - What automated and manual tests prove readiness? - Who reviewed the evidence and approved production release? - What were the issues that were reported - and what changed between versions to resolve them? - What is happening in production, and which failures require human review? - How does production feedback become new evaluation data? ## Databricks AI Gateway value prop Databricks describes Unity AI Gateway as its governance solution for enterprise AI, built on Unity Catalog. Its core value is controlling and observing AI service interactions in Databricks: - **Access governance:** Register models, agents, MCP services, functions, and HTTP connections as Unity Catalog securable objects, then grant or revoke access with Unity Catalog privileges. - **Traffic management:** Route requests across model services and MCP services, configure traffic splitting and fallbacks, apply rate limits, and manage budgets with hard spend caps. - **Runtime guardrails:** Attach service policies to allow, deny, or require approval for interactions based on request and response content. Databricks lists built-in guardrails for PII, prompt injection, and unsafe content. - **Usage, cost, and audit logging:** Track requests, token usage, latency, requester identity, tags, model destination, and cost through system tables, dashboards, and Unity Catalog Delta inference tables. - **Databricks-native integration:** Keep governance close to Databricks-hosted models, external models exposed through model services, AI agents, MCP services, and Unity Catalog. References: - [Databricks: AI governance with Unity AI Gateway](https://docs.databricks.com/aws/en/ai-gateway/) - [Databricks: Usage tracking](https://docs.databricks.com/aws/en/ai-gateway/usage-tracking) - [Databricks: Rate limits](https://docs.databricks.com/aws/en/ai-gateway/rate-limits) - [Databricks: Inference tables](https://docs.databricks.com/aws/en/ai-gateway/inference-tables) ## GGX value add across automated compliance dimensions GGX complements a gateway by turning runtime controls and logs into a governed compliance process. | Compliance dimension | Databricks AI Gateway contribution | GGX value add | | ------------------------------ | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | AI inventory and ownership | AI services can be governed as Unity Catalog securable objects. | GGX maintains a broader GenAI inventory for models, prompts, RAGs, pipelines, tables, reports, global functions, versions, owners, access scope, and review status. | | Versioning and change control | Gateway configuration and Unity Catalog permissions govern current service access. | GGX tracks version history, object lineage, challenger comparisons, approvals, and locked production artifacts so reviewers can see what changed and why it was accepted. | | Automated risk identification | Built-in service policies can detect or block selected runtime risks such as PII, prompt injection, and unsafe content. | GGX maps use-case-specific risks across accuracy, stability, bias, toxicity, vulnerability, privacy leakage, groundedness, hallucination, dark patterns, and agent tool behavior. | | Pre-production validation | Gateway can control which AI services are available and enforce runtime policies. | GGX runs reproducible simulations and comparison jobs against curated datasets, expected outputs, thresholds, and standardized reports before launch. | | Model risk management evidence | Databricks tables can show usage, requester, tokens, latency, payloads, and policy outcomes. | GGX packages evaluation results into dashboards and approval evidence aligned to Model Risk Management, Fair Lending, technology, infosec, legal, and business review workflows. | | Fairness and bias testing | Gateway policies can block or flag selected content patterns at runtime. | GGX supports explicit bias and fairness testing through reusable reports, curated datasets, segment analysis, dashboards, and approval-ready outputs. | | Human oversight | Databricks service policies can require approval for selected interactions. | GGX supports human testing, feedback portals, findings workflows, annotation queues, and business SME review loops that turn expert judgment into reusable ground truth. | | Production monitoring | Gateway usage and inference tables provide operational telemetry and request/response logs. | GGX ingests production data, runs monitoring reports, alerts on threshold breaches, routes records for annotation, and turns production findings into new test cases. | | Auditability and documentation | Databricks provides system tables, dashboards, and inference tables inside Unity Catalog. | GGX links tests, results, approvals, findings, versions, lineage, and monitoring evidence into an auditable lifecycle record that can be exported or reviewed by committees. | | Ecosystem coverage | Strongest for AI services governed through Databricks and Unity Catalog. | GGX can govern AI applications across multiple providers, agent frameworks, RAG systems, custom pipelines, and production environments, including Databricks where it is part of the stack. | ## When to use both together Use Databricks AI Gateway and GGX together when Databricks is part of the production AI stack and the organization needs a defensible approval process. 1. **Develop and serve in Databricks:** Use Databricks model services, Unity Catalog permissions, service policies, budgets, rate limits, usage tables, and inference tables. 2. **Register in GGX:** Register the AI application or pipeline, plus relevant models, prompts, RAGs, datasets, and reports. 3. **Evaluate before launch:** Run GGX simulations and comparison jobs for business quality and compliance dimensions. 4. **Approve with evidence:** Attach results to GGX approval workflows for business, risk, compliance, and technical sign-off. 5. **Deploy in databricks:** Export the agent from GGX and deploy it into databricks. 5. **Monitor after launch:** Feed Databricks inference logs or production traces into GGX monitoring dashboards and annotation queues. 6. **Close the loop:** Convert production findings into new test cases, compare challenger versions, and maintain a complete lifecycle record. ## Related GGX docs - [Inventory Management](../../register-and-refine/inventory-management/) - [Lineage Tracking](../../register-and-refine/lineage-tracking/) - [Evaluations and Approval](../../evaluate-and-approve/) - [Reporting](../../evaluate-and-approve/reporting/) - [Approval Workflows](../../evaluate-and-approve/approval-workflows/) - [Deployment and Monitoring](../../deploy-and-monitor/) - [Annotation Queues](../../deploy-and-monitor/annotation-queues/) --- # LiteLLM Gateway Source: https://docs.genguardx.ai/integrations/llm-gateways/litellm/ Markdown: https://docs.genguardx.ai/integrations/llm-gateways/litellm/index.md Description: Register LiteLLM as a GGX Model Registry model and route GGX requests to any LLM provider configured behind LiteLLM. [LiteLLM](https://docs.litellm.ai/docs/) provides a unified interface for many LLM providers and includes a self-hosted proxy server that can expose an OpenAI-compatible gateway. LiteLLM is commonly used to standardize provider APIs, manage virtual keys, track spend, configure fallbacks, and route requests across model deployments. In GGX, LiteLLM can be registered in the [Model Registry](../../register-and-refine/inventory-management/model-catalog/) as a Model. The registered GGX Model calls your LiteLLM proxy, and LiteLLM can route the request to any underlying LLM that your LiteLLM configuration exposes. ## When to use this integration Use LiteLLM with GGX when: - Your engineering team already uses LiteLLM as the LLM proxy. - You want GGX pipelines to call the same model aliases that production applications call. - You want to evaluate several providers through one gateway layer. - You want LiteLLM to handle provider credentials, routing, fallback, or cost tracking while GGX handles evaluation, approval, and monitoring. ## Register LiteLLM in the Model Registry Create one GGX Model per LiteLLM model alias, or create one parameterized model that accepts a `model` argument. | GGX setting | Recommended value | | --- | --- | | **Name** | `litellm_` or `litellm_gateway` | | **Description** | Note the LiteLLM proxy URL, allowed model aliases, provider list, and fallback policy. | | **Model Provider** | Use a custom/API-based model or Python function that calls the LiteLLM proxy. | | **Arguments** | `messages`, `model`, `temperature`, `max_tokens` | | **Environment variables** | `LITELLM_API_KEY`, `LITELLM_BASE_URL`, `LITELLM_MODEL` | ## Example scoring logic ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("LITELLM_API_KEY"), base_url=os.getenv("LITELLM_BASE_URL"), ) selected_model = model if model else os.getenv("LITELLM_MODEL") completion = client.chat.completions.create( model=selected_model, messages=messages, temperature=float(temperature), max_tokens=int(max_tokens), ) return { "output": completion.choices[0].message.content, "model": selected_model, } ``` ## Governance notes - Register each approved LiteLLM model alias as a separate GGX Model if business or risk reviewers need model-level approval. - Document whether LiteLLM fallback is enabled, because fallback may cause a single GGX Model to use more than one underlying provider. - Include LiteLLM request IDs, selected model names, or provider metadata in GGX outputs when available, so evaluation and monitoring dashboards can segment results by actual route. Reference: [LiteLLM documentation](https://docs.litellm.ai/docs/) --- # OpenRouter Gateway Source: https://docs.genguardx.ai/integrations/llm-gateways/openrouter/ Markdown: https://docs.genguardx.ai/integrations/llm-gateways/openrouter/index.md Description: Register OpenRouter as a GGX Model Registry model and route GGX requests to OpenRouter-supported models through a unified API. [OpenRouter](https://openrouter.ai/docs/quickstart) provides a unified API for accessing many AI models through a single endpoint. It supports direct API calls and OpenAI-compatible SDK usage, so GGX can call OpenRouter as a gateway-backed model. In GGX, OpenRouter can be registered in the [Model Registry](../../register-and-refine/inventory-management/model-catalog/) as a Model. The registered GGX Model calls OpenRouter, and OpenRouter can route to any model that your OpenRouter account and request configuration can access. ## When to use this integration Use OpenRouter with GGX when: - You want a single model-access endpoint for many third-party models. - You want to compare models available through OpenRouter using GGX simulations and comparison jobs. - You want GGX pipelines to use OpenRouter model slugs instead of provider-specific SDKs. - You need quick access to multiple hosted models while keeping GGX as the governance and approval layer. ## Register OpenRouter in the Model Registry | GGX setting | Recommended value | | --- | --- | | **Name** | `openrouter_` or `openrouter_gateway` | | **Description** | Identify the OpenRouter model slug, routing preferences, and whether fallback or model aliases are used. | | **Model Provider** | Use a custom/API-based model or Python function that calls OpenRouter. | | **Arguments** | `messages`, `model`, `temperature`, `max_tokens` | | **Environment variables** | `OPENROUTER_API_KEY`, `OPENROUTER_MODEL` | ## Example scoring logic ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("OPENROUTER_API_KEY"), base_url="https://openrouter.ai/api/v1", ) selected_model = model if model else os.getenv("OPENROUTER_MODEL") completion = client.chat.completions.create( model=selected_model, messages=messages, temperature=float(temperature), max_tokens=int(max_tokens), ) return { "output": completion.choices[0].message.content, "model": selected_model, } ``` ## Governance notes - Register one GGX Model per OpenRouter model slug when each model needs separate approval evidence. - Use GGX comparison jobs to evaluate challenger OpenRouter models against the same dataset and reports. - Record the OpenRouter model slug in the GGX Model description and exported evidence. - If you use OpenRouter aliases or routing controls, document them in the GGX Risk Assessment so reviewers understand what model may answer a request. Reference: [OpenRouter quickstart](https://openrouter.ai/docs/quickstart) --- # Portkey Gateway Source: https://docs.genguardx.ai/integrations/llm-gateways/portkey/ Markdown: https://docs.genguardx.ai/integrations/llm-gateways/portkey/index.md Description: Register Portkey as a GGX Model Registry model and use Portkey provider routing with GGX evaluations, pipelines, approvals, and monitoring. [Portkey](https://docs.portkey.ai/docs/introduction/what-is-portkey) is an AI gateway that provides a unified interface for many AI models, with tooling for visibility, routing, control, and security. Portkey can be used directly through its SDK or through OpenAI-compatible clients pointed at the Portkey gateway. In GGX, Portkey can be registered in the [Model Registry](../../register-and-refine/inventory-management/model-catalog/) as a Model. The registered GGX Model calls Portkey, and Portkey can connect to any LLM provider or model that your Portkey configuration allows. ## When to use this integration Use Portkey with GGX when: - Portkey is the enterprise gateway for model access. - Teams use Portkey provider configs, routing rules, or observability in production. - You want GGX evaluation and approval evidence to exercise the same gateway route as production. - You need GGX to test multiple Portkey-backed providers without adding each provider directly to GGX. ## Register Portkey in the Model Registry | GGX setting | Recommended value | | --- | --- | | **Name** | `portkey_` or `portkey_gateway` | | **Description** | Identify the Portkey provider, config, model alias, routing behavior, and any data logging policy. | | **Model Provider** | Use a custom/API-based model or Python function that calls the Portkey gateway. | | **Arguments** | `messages`, `model`, `temperature`, `max_tokens`, optional `provider` or config identifier | | **Environment variables** | `PORTKEY_API_KEY`, `PORTKEY_BASE_URL`, `PORTKEY_PROVIDER`, `PORTKEY_MODEL` | ## Example scoring logic ```python import os from openai import OpenAI client = OpenAI( api_key=os.getenv("PORTKEY_UPSTREAM_API_KEY", "not-used"), base_url=os.getenv("PORTKEY_BASE_URL"), default_headers={ "x-portkey-api-key": os.getenv("PORTKEY_API_KEY"), "x-portkey-provider": provider if provider else os.getenv("PORTKEY_PROVIDER"), }, ) selected_model = model if model else os.getenv("PORTKEY_MODEL") completion = client.chat.completions.create( model=selected_model, messages=messages, temperature=float(temperature), max_tokens=int(max_tokens), ) return { "output": completion.choices[0].message.content, "model": selected_model, } ``` ## Governance notes - Register separate GGX Models for Portkey providers or configs that require separate approvals. - Capture the Portkey provider/config name in the GGX Model description and Risk Assessment. - If Portkey routing can shift between upstream LLMs, include route metadata in GGX evaluation outputs whenever available. - Use GGX simulations and comparisons to test whether Portkey routing changes affect accuracy, stability, safety, latency, or cost-sensitive behavior. Reference: [Portkey documentation](https://docs.portkey.ai/docs/introduction/what-is-portkey) --- # Setting up Integrations Source: https://docs.genguardx.ai/integrations/llm-providers/ Markdown: https://docs.genguardx.ai/integrations/llm-providers/index.md Description: Configure LLM provider integrations in GGX by adding API keys, testing connections, and exposing secure environment variables for model registration and model code. Setup integrations to services that provide LLMs as an API. Configure your API keys once to access multiple AI providers across the platform. ## Getting Started Navigate to **Settings > Platform Integrations** to configure your LLM providers. Each provider requires an API key that creates secure environment variables for your models. ## LLM Providers ### OpenAI **Models available**: o4-mini, o3, o3-mini, etc 1. Click the OpenAI card 2. Enter your API key from platform.openai.com 3. Test connection and save ### Anthropic **Models available**: claude-4-opus, claude-4-sonnet, claude-3.7-sonnet, etc 1. Click the Anthropic card 2. Enter your API key from console.anthropic.com 3. Test connection and save ### Azure AI **Models available**: o4-mini, o3, o3-mini, etc 1. Click the Azure AI card 2. Enter your Azure OpenAI API key 3. Test connection and save ### Amazon Bedrock **Models available**: amazon.titan-text-premier-v1:0, amazon.titan-text-express-v1, amazon.titan-text-lite-v1, etc 1. Click the Amazon Bedrock card 2. Configure AWS credentials: * **Access Key ID**: Your AWS access key * **Secret Access Key**: Your AWS secret key * **Session Token**: Optional temporary session token * **Default Region**: AWS region (e.g., us-east-1) 3. Test connection and save ### DeepSeek **Models available**: deepseek-r1, deepseek-v3, deepseek-v2.5 1. Click the DeepSeek card 2. Get your API key from DeepSeek Platform 3. Enter your API key in the field 4. Click "Test Connection" to verify 5. Click "Save" to complete setup ### Google Vertex AI **Models available**: gemini-2.5-pro, gemini-2.5-flash, gemini-2.0-flash, etc 1. Click the Google Vertex AI card 2. Upload service account JSON key 3. Test connection and save ### Hugging Face **Models available**: llama-3.1-8b-instruct, llama-3.1-70b-instruct, llama-3.1-405b-instruct, etc 1. Click the Hugging Face card 2. Enter your HF token from huggingface.co/settings/tokens 3. Test connection and save ### Nvidia NIM **Models available**: llama-3.1-nemotron-instruct-70b, llama-3.3-nemotron-super-49b-reasoning, llama-3.1-nemotron-ultra-253b-v1-reasoning, etc 1. Click the Nvidia NIM card 2. Enter your Nvidia API key 3. Test connection and save ### GitHub Models **Models available**: o4-mini, o3, o3-mini, etc 1. Click the GitHub Models card 2. Enter your GitHub token 3. Test connection and save ## Integration Status * **Active**: Ready to use in model registration * **Inactive**: Needs configuration or has connection issues ## Setting Up API Keys Each integration creates environment variables that you can use in your model code: * **OpenAI**: `OPENAI_API_KEY` * **Anthropic**: `ANTHROPIC_API_KEY` * **Azure AI**: `AZURE_ENDPOINT` * **DeepSeek**: `DEEPSEEK_API_KEY` * **Google Vertex AI**: `GOOGLE_API_TOKEN` * **Hugging Face**: `HUGGING_FACE_HUB_TOKEN` * **GitHub Models**: `GITHUB_TOKEN` * **Amazon Bedrock**: `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN`, `AWS_DEFAULT_REGION` ## Using Environment Variables Once configured, the environment variables are automatically available in your model code: ```python import os # Access your API keys openai_key = os.getenv("OPENAI_API_KEY") anthropic_key = os.getenv("ANTHROPIC_API_KEY") azure_endpoint = os.getenv("AZURE_ENDPOINT") deepseek_key = os.getenv("DEEPSEEK_API_KEY") google_token = os.getenv("GOOGLE_API_TOKEN") hf_token = os.getenv("HUGGING_FACE_HUB_TOKEN") github_token = os.getenv("GITHUB_TOKEN") # AWS Bedrock credentials aws_access_key = os.getenv("AWS_ACCESS_KEY_ID") aws_secret_key = os.getenv("AWS_SECRET_ACCESS_KEY") aws_session_token = os.getenv("AWS_SESSION_TOKEN") aws_region = os.getenv("AWS_DEFAULT_REGION") ``` ## Need Help? * Use "Test Connection" to verify your API keys * Check provider documentation for API key setup * Ensure proper permissions for your API keys --- # Anthropic Integration Source: https://docs.genguardx.ai/integrations/llm-providers/anthropic/ Markdown: https://docs.genguardx.ai/integrations/llm-providers/anthropic/index.md Description: Configure Anthropic in GGX and register Claude-backed models with provider credentials, model arguments, scoring logic, and governed model metadata. The Anthropic integration provides access to Claude models through a unified interface. Configure once, use everywhere with enterprise-grade safety features and constitutional AI principles built into every interaction. ## Integrating Anthropic Simply enter your Anthropic API key once in the Platform Integrations section. This enables authorized users to access Claude models within the platform. Once integrated, models can be registered and used as any other python object on the platform. ```python # Example: Using a registered Anthropic model result = claude_sonnet_model(text="Analyze this document", max_tokens=1500) ``` ## Example of Models Supported Anthropic provides access to Claude models with different capabilities and performance characteristics: **Claude 3.5 Sonnet** - Balanced performance for most use cases with strong reasoning **Claude 3 Opus** - Most capable model for complex tasks requiring deep analysis **Claude 3 Haiku** - Fast, cost-effective model for simple tasks **Claude 3.5 Haiku** - Enhanced version with improved speed and capabilities **Additional Models** - Latest Claude variants with enhanced reasoning capabilities ## Registering a New Anthropic Model Navigate to **New Model** to begin registration. The registration form connects your Anthropic integration with custom model configurations. ### Basic Information **Description**: Document your model's purpose, use cases, and limitations. For example: "Claude 3.5 Sonnet optimized for document analysis and content generation. Use for enterprise content processing with built-in safety features. Ideal for complex reasoning tasks and multi-step analysis." ### Code Configuration **Alias**: A unique identifier for your model (e.g., `claude_sonnet_analyzer`, `claude_document_processor`). This becomes the variable name you'll use in code. **Output Type**: Define the return format: - `Map[String, String]` - Key-value pairs for structured responses - `String` - Simple text responses - `List` - Array of items **Input Type**: Select your implementation approach: - **API Based**: Platform handles API calls automatically using your Anthropic integration - **Python Function**: Custom function implementation with full control - **Custom**: Advanced configurations for specialized use cases **Model Provider**: Select "Anthropic" from your configured integrations. ### Arguments Configuration Define input parameters that your model will accept. Important: Variables declared here are automatically available in the Scoring Logic section. Common argument patterns for Anthropic models: | Alias | Type | Optional | Default Value | Usage | |-------|------|----------|---------------|-------| | `text` | String | No | N/A | Main input content | | `max_tokens` | Numerical | Yes | 1500 | Maximum response length | | `temperature` | Numerical | Yes | 0.7 | Controls response creativity | | `system_prompt` | String | Yes | "" | System instructions | Use **+ Add Argument** to include additional parameters. ### Scoring Logic Implementation In the Scoring Logic section, you can directly reference any variable declared in the Arguments section. The platform automatically makes these available in your code. Example implementation for a text analysis model: ```python # Arguments: text, max_tokens, temperature are automatically available import os import anthropic # Direct initialization client = anthropic.Anthropic( api_key=os.getenv("ANTHROPIC_API_KEY") ) if text is None: return None message = client.messages.create( model="claude-3-5-sonnet-20241022", max_tokens=int(max_tokens), temperature=float(temperature), system=system_prompt if system_prompt else "You are a helpful AI assistant.", messages=[ {"role": "user", "content": text} ] ) return {"output": message.content[0].text, "context": None} ``` ## Platform Integration Setup Before registering models, configure your Anthropic credentials: 1. Navigate to **Settings > Platform Integrations** 2. Click on **Anthropic** 3. Enter your Anthropic API key 4. Test the connection The platform creates environment variables automatically: - `ANTHROPIC_API_KEY` ## Example Use Case: Document Analysis Model An Anthropic Claude model configured for enterprise document analysis demonstrates the complete workflow: ### Arguments Configuration: - `text` (String, required) - `max_tokens` (Numerical, optional, default: "2000") - `temperature` (Numerical, optional, default: "0.3") - `system_prompt` (String, optional, default: "You are a helpful document analysis assistant.") ### Usage: ```python # Model becomes available as: claude_document_analyzer result = claude_document_analyzer( text="Your document text here...", max_tokens=2000, temperature=0.3, system_prompt="Analyze this document and provide key insights with detailed explanations." ) ``` ## Want to Learn More? - Review [Anthropic documentation](https://docs.anthropic.com/) - Check [Claude model capabilities](https://www.anthropic.com/claude) - Monitor usage through Anthropic Console - Set up rate limiting and cost management --- # AWS Bedrock Integration Source: https://docs.genguardx.ai/integrations/llm-providers/aws-bedrock/ Markdown: https://docs.genguardx.ai/integrations/llm-providers/aws-bedrock/index.md Description: Configure AWS Bedrock in GGX and register Bedrock-backed foundation models with AWS credentials, provider settings, model arguments, and scoring logic. The AWS Bedrock integration provides access to foundation models from multiple AI providers through a unified interface. Configure once, use everywhere with enterprise-grade security, scalability, and comprehensive model selection from leading AI companies. ## Integrating AWS Bedrock Configure your AWS credentials once in the Platform Integrations section. This enables authorized users to access Bedrock foundation models within the platform. Once integrated, models can be registered and used as any other python object on the platform. ```python # Example: Using a registered Bedrock model result = bedrock_claude_model(text="Analyze this data", temperature=0.8) ``` ## Example of Models Supported AWS Bedrock provides access to foundation models from multiple leading AI providers: **Amazon Titan** - Amazon's own foundation models for text generation and embeddings **Anthropic Claude** - Claude 3.7 Sonnet, Claude 3.5 Sonnet, and other Claude variants **Meta Llama** - Llama 3.2 series with fine-tuning capabilities **Cohere Command** - Command R/R+ and Embed v3 families for text generation and embeddings **AI21 Labs Jamba** - Jamba 1.5 series for advanced language processing **Stability AI** - Stable Diffusion series for image generation **Additional Models** - +122 models available through Amazon Bedrock Marketplace ## Registering a New Bedrock Model Navigate to **New Model** to begin registration. The registration form connects your AWS Bedrock integration with custom model configurations. ### Basic Information **Description**: Document your model's purpose, use cases, and limitations. For example: "Claude 3.7 Sonnet on AWS Bedrock optimized for enterprise content analysis. Use for document processing with AWS compliance features. Ideal for complex reasoning and multi-step analysis tasks." ### Code Configuration **Alias**: A unique identifier for your model (e.g., `bedrock_claude_analyzer`, `titan_embedder`). This becomes the variable name you'll use in code. **Output Type**: Define the return format: - `Map[String, String]` - Key-value pairs for structured responses - `String` - Simple text responses - `List` - Array of items **Input Type**: Select your implementation approach: - **API Based**: Platform handles API calls automatically using your AWS integration - **Python Function**: Custom function implementation with full control - **Custom**: Advanced configurations for specialized use cases **Model Provider**: Select "Amazon Bedrock" from your configured integrations. ### Arguments Configuration Define input parameters that your model will accept. Important: Variables declared here are automatically available in the Scoring Logic section. Common argument patterns for Bedrock models: | Alias | Type | Optional | Default Value | Usage | |-------|------|----------|---------------|-------| | `text` | String | No | N/A | Main input content | | `temperature` | Numerical | Yes | 0.7 | Controls response creativity | | `max_tokens` | Numerical | Yes | 1500 | Maximum response length | | `system_prompt` | String | Yes | "" | System instructions | Use **+ Add Argument** to include additional parameters. ### Scoring Logic Implementation In the Scoring Logic section, you can directly reference any variable declared in the Arguments section. The platform automatically makes these available in your code. ```python # Arguments: text, temperature are automatically available import os import boto3 # Direct initialization client = boto3.client( service_name="bedrock-runtime", region_name=os.getenv("AWS_DEFAULT_REGION", "us-east-1"), aws_access_key_id=os.getenv("AWS_ACCESS_KEY_ID"), aws_secret_access_key=os.getenv("AWS_SECRET_ACCESS_KEY"), aws_session_token=os.getenv("AWS_SESSION_TOKEN") ) if text is None: return None # Prepare the conversation for Claude models conversation = [ { "role": "user", "content": [{"text": text}] } ] # Add system message if provided if system_prompt: conversation.insert(0, { "role": "system", "content": [{"text": system_prompt}] }) response = client.converse( modelId="anthropic.claude-3-5-sonnet-20240620-v1:0", messages=conversation, inferenceConfig={ "temperature": float(temperature), "maxTokens": int(max_tokens) } ) return { "output": response["output"]["message"]["content"][0]["text"], "context": response["usage"]["inputTokens"] + response["usage"]["outputTokens"] } ``` ## Platform Integration Setup Before registering models, configure your AWS credentials: 1. Navigate to **Settings > Platform Integrations** 2. Click on **Amazon Bedrock** 3. Configure AWS credentials: * **Access Key ID**: Your AWS access key * **Secret Access Key**: Your AWS secret key * **Session Token**: Optional temporary session token * **Default Region**: AWS region (e.g., us-east-1) 4. Test the connection The platform creates environment variables automatically: - `AWS_ACCESS_KEY_ID` - `AWS_SECRET_ACCESS_KEY` - `AWS_SESSION_TOKEN` - `AWS_DEFAULT_REGION` ## Example Use Case: Document Analysis Model An AWS Bedrock Claude model configured for enterprise document analysis demonstrates the complete workflow: ### Arguments Configuration: - `text` (String, required) - `temperature` (Numerical, optional, default: "0.3") - `max_tokens` (Numerical, optional, default: "2000") - `system_prompt` (String, optional, default: "You are an expert document analyst.") ### Usage: ```python # Model becomes available as: document_analyzer result = document_analyzer( text="Your document content here...", temperature=0.3, max_tokens=2000, system_prompt="Analyze this document for key insights, themes, and actionable recommendations." ) ``` ## Want to Learn More? - Review [AWS Bedrock documentation](https://docs.aws.amazon.com/bedrock/) - Check [supported foundation models](https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html) - Monitor usage through AWS Console - Set up cost management and billing alerts --- # Azure AI Integration Source: https://docs.genguardx.ai/integrations/llm-providers/azureai/ Markdown: https://docs.genguardx.ai/integrations/llm-providers/azureai/index.md Description: Configure Azure AI in GGX and register Azure-hosted OpenAI models with API credentials, model provider settings, arguments, and scoring logic. The Azure AI integration provides access to OpenAI models hosted on Microsoft Azure through a unified interface. Configure once, use everywhere with enterprise-grade security and compliance. ## Integrating Azure AI Simply enter your Azure OpenAI API key and endpoint once in the Platform Integrations section. This enables authorized users to access Azure-hosted OpenAI models within the platform. Once integrated, models can be registered and used as any other python object on the platform. ```python # Example: Using a registered Azure model result = azure_gpt4_model(text="Analyze this data", temperature=0.8) ``` ## Supported Models Azure AI provides access to OpenAI models hosted on Microsoft Azure: **GPT-4** - Advanced reasoning and complex task completion **GPT-3.5 Turbo** - Fast, efficient responses for most use cases **o4-mini, o3, o3-mini** - Latest OpenAI models with enhanced capabilities **Additional Models** - other OpenAI variants available ## Registering a New Azure Model Navigate to **New Model** to begin registration. The registration form connects your Azure AI integration with custom model configurations. ### Basic Information **Description**: Document your model's purpose, use cases, and limitations. For example: "GPT-4 on Azure optimized for document analysis. Use for enterprise content processing with Azure compliance. Ideal for sensitive data workflows." ### Code Configuration **Alias**: A unique identifier for your model (e.g., `azure_gpt4_analyzer`, `ggx_gpt4`). This becomes the variable name you'll use in code. **Output Type**: Define the return format: - `Map[String, String]` - Key-value pairs for structured responses - `String` - Simple text responses - `List` - Array of items **Input Type**: Select your implementation approach: - **API Based**: Platform handles API calls automatically using your Azure integration - **Python Function**: Custom function implementation with full control - **Custom**: Advanced configurations for specialized use cases **Model Provider**: Select "Azure AI" from your configured integrations. ### Arguments Configuration Define input parameters that your model will accept. Important: Variables declared here are automatically available in the Scoring Logic section. Common argument patterns for Azure models: | Alias | Type | Optional | Default Value | Usage | |-------|------|----------|---------------|-------| | `text` | String | No | N/A | Main input content | | `temperature` | Numerical | Yes | 0.7 | Controls response creativity | | `max_tokens` | Numerical | Yes | 1500 | Maximum response length | | `system_prompt` | String | Yes | "" | System instructions | Use **+ Add Argument** to include additional parameters. ### Scoring Logic Implementation In the Scoring Logic section, you can directly reference any variable declared in the Arguments section. The platform automatically makes these available in your code. ```python # Arguments: text, temperature are automatically available import os from openai import AzureOpenAI # Direct initialization client = AzureOpenAI( azure_endpoint="https://example-openai-resource.openai.azure.com/", api_key=os.getenv("AZURE_OPENAI_API_KEY"), api_version="2024-12-01-preview", ) chat_prompt = [{"role": "system", "content": [{"type": "text", "text": text}]}] completion = client.chat.completions.create( model="gpt-4.1", messages=chat_prompt, max_tokens=1500, temperature=float(temperature), top_p=0.95, frequency_penalty=0, presence_penalty=0, stop=None, stream=False, ) return {"output": completion.choices[0].message.content, "context": None} ``` ## Platform Integration Setup Before registering models, configure your Azure credentials: 1. Navigate to **Settings > Platform Integrations** 2. Click on **Azure AI** 3. Enter your Azure OpenAI API key 4. Provide your Azure endpoint URL 5. Test the connection The platform creates environment variables automatically: - `AZURE_ENDPOINT` ## Example Use Case: Document Processing Model An Azure GPT-4 model configured for enterprise document processing demonstrates the complete workflow: ### Arguments Configuration: - `text` (String, required) - `temperature` (Numerical, optional, default: "0.3") - `max_tokens` (Numerical, optional, default: "1500") - `system_prompt` (String, optional, default: "") ### Usage: ```python # Model becomes available as: document_processor result = document_processor( text="Your document text here...", temperature=0.3, max_tokens=2000, system_prompt="Process the document and extract key information." ) ``` ## Want to Learn More? - Review Azure OpenAI documentation - Check Azure compliance and security features - Monitor usage through Azure Portal - Set up cost management and billing alerts --- # Google Cloud Vertex AI Integration Source: https://docs.genguardx.ai/integrations/llm-providers/gcp-vertexai/ Markdown: https://docs.genguardx.ai/integrations/llm-providers/gcp-vertexai/index.md Description: Configure Google Cloud Vertex AI in GGX and register Gemini models with service credentials, provider settings, input arguments, and scoring logic. The Google Cloud Vertex AI integration provides access to Gemini models and AI agents through a unified interface. Configure once, use everywhere with enterprise-grade security and scalability. ## Integrating Google Cloud Vertex AI Simply upload your service account JSON key once in the Platform Integrations section. This enables authorized users to access Google Cloud Vertex AI models within the platform. Once integrated, models can be registered and used as any other python object on the platform. ```python # Example: Using a registered Gemini model result = gemini_model(text="Analyze this data", temperature=0.8) ``` ## Supported Models Google Cloud Vertex AI provides access to Gemini models and agents: **Gemini 2.5 Pro** - Advanced reasoning and multimodal capabilities **Gemini 2.5 Flash** - Fast, efficient responses for high-volume use cases **Gemini 2.0 Flash** - Latest generation model with improved performance **Additional Models** - other Gemini variants available ## Registering a New Gemini Model Navigate to **New Model** to begin registration. The registration form connects your Google Cloud integration with custom model configurations. ### Basic Information **Description**: Document your model's purpose, use cases, and limitations. For example: "Gemini 2.5 Pro optimized for content analysis. Use for document summarization and multimodal tasks. Ideal for complex reasoning workflows." ### Code Configuration **Alias**: A unique identifier for your model (e.g., `gemini_analyzer`, `content_summarizer`). This becomes the variable name you'll use in code. **Output Type**: Define the return format: - `Map[String, String]` - Key-value pairs for structured responses - `String` - Simple text responses - `List` - Array of items **Input Type**: Select your implementation approach: - **API Based**: Platform handles API calls automatically using your Google Cloud integration - **Python Function**: Custom function implementation with full control - **Custom**: Advanced configurations for specialized use cases **Model Provider**: Select "Google Vertex AI" from your configured integrations. ### Arguments Configuration Define input parameters that your model will accept. Important: Variables declared here are automatically available in the Scoring Logic section. Common argument patterns for Gemini models: | Alias | Type | Optional | Default Value | Usage | |-------|------|----------|---------------|-------| | `text` | String | No | N/A | Main input content | | `temperature` | Numerical | Yes | 0.7 | Controls response creativity | | `system_instruction` | String | Yes | "" | System prompt for model behavior | | `seed` | Numerical | Yes | 2025 | Deterministic generation seed | Use **+ Add Argument** to include additional parameters. ### Scoring Logic Implementation In the Scoring Logic section, you can directly reference any variable declared in the Arguments section. The platform automatically makes these available in your code. ```python # Arguments: text, temperature, system_instruction are automatically available import os from google import genai from google.genai import types client = genai.Client(api_key=os.getenv("GOOGLE_API_TOKEN")) config = types.GenerateContentConfig( temperature=temperature, seed=2025, system_instruction=system_instruction ) response = client.models.generate_content( model="gemini-2.0-flash", contents=text, config=config ) return {"response": response.text} ``` ## Platform Integration Setup Before registering models, configure your Google Cloud credentials: 1. Navigate to **Settings > Platform Integrations** 2. Click on **Google Vertex AI** 3. Upload your service account JSON key file 4. Enter your Google Cloud project ID 5. Test the connection The platform creates environment variables automatically: - `GOOGLE_API_TOKEN` - `GOOGLE_CLOUD_PROJECT` ## Example Use Case: Content Analysis Model A Gemini 2.5 Pro model configured for content analysis demonstrates the complete workflow: ### Arguments Configuration: - `text` (String, required) - `temperature` (Numerical, optional, default: "0.3") - `system_instruction` (String, optional, default: "You are an expert content analyst.") - `seed` (Numerical, optional, default: "2025") ### Usage: ```python # Model becomes available as: content_analyzer result = content_analyzer( text="Your document content here...", temperature=0.3, system_instruction="Provide a comprehensive analysis with key insights and recommendations.", seed=2025 ) ``` ## Want to Learn More? - Review Google Cloud Vertex AI documentation - Check Gemini model specifications - Monitor usage and costs through Google Cloud Console - Set up billing alerts for cost management --- # Hugging Face Integration Source: https://docs.genguardx.ai/integrations/llm-providers/huggingface/ Markdown: https://docs.genguardx.ai/integrations/llm-providers/huggingface/index.md Description: Register Hugging Face models in GGX using Hub models, optional tokens, model metadata, input arguments, and custom Python scoring logic. The Hugging Face integration provides direct access to thousands of open-source models from the Hugging Face Hub through a unified interface. Load once, cache efficiently, and use enterprise-grade model management with full governance and tracking. ## Integrating Hugging Face No platform-level integration required - Hugging Face models are accessed directly through the `transformers` library. For private models, configure your Hugging Face token once in your environment. Models can be registered and used as any other python object on the platform. ```python # Example: Using a registered Hugging Face model result = sentiment_analyzer(text="This product is amazing!", confidence_threshold=0.8) ``` ## Supported Models Hugging Face provides access to thousands of models across multiple categories: **Text Classification** - Sentiment analysis, content moderation, topic classification **Text Generation** - LLaMA, Mistral, CodeLlama, and other foundation models **Text Embedding** - Sentence transformers, multilingual embeddings **Guardrail Models** - Prompt injection detection, content safety filters **Additional Models** - other models available on Hugging Face Hub ## Registering a New Hugging Face Model Navigate to **New Model** to begin registration. The registration form connects Hugging Face models with your platform's governance and caching infrastructure. ### Basic Information **Description**: Document your model's purpose, use cases, and limitations. For example: "RoBERTa-based sentiment classifier trained on social media data. Use for customer feedback analysis and content moderation. Optimized for short text under 512 tokens." ### Code Configuration **Alias**: A unique identifier for your model (e.g., `sentiment_classifier`, `prompt_guard`, `sentence_embedder`). This becomes the variable name you'll use in code. **Output Type**: Define the return format: - `Map[String, String]` - Key-value pairs for structured responses - `String` - Simple text responses - `List` - Array of items **Input Type**: Select your implementation approach: - **API Based**: Platform handles API calls automatically using your integration - **Python Function**: Custom function implementation with full control - **Custom**: Advanced configurations for specialized use cases **Model Provider**: Select "Hugging Face" from your configured integrations. ### Arguments Configuration Define input parameters that your model will accept. Important: Variables declared here are automatically available in the Scoring Logic section. Common argument patterns for Hugging Face models: | Alias | Type | Optional | Default Value | Usage | |-------|------|----------|---------------|-------| | `text` | String | No | N/A | Main input content | | `max_length` | Numerical | Yes | 512 | Maximum token length | | `temperature` | Numerical | Yes | 0.7 | Generation randomness | | `threshold` | Numerical | Yes | 0.5 | Classification threshold | Use **+ Add Argument** to include additional parameters. ### Scoring Logic Implementation In the Scoring Logic section, you can directly reference any variable declared in the Arguments section. The platform automatically makes these available in your code. Example implementation for a text classification model: ```python # Arguments: text, threshold are automatically available import os from transformers import pipeline from huggingface_hub import login # Authentication for private models (optional) login(token=os.getenv("HUGGINGFACE_TOKEN")) # Direct initialization client = pipeline( "text-classification", model="cardiffnlp/twitter-roberta-base-sentiment-latest" ) if text is None: return None results = client(text) confidence = results[0]["score"] if confidence >= threshold: return {"sentiment": results[0]["label"], "confidence": confidence} else: return {"sentiment": "uncertain", "confidence": confidence} ``` ## Platform Integration Setup For private models, configure your Hugging Face credentials: 1. Navigate to **Settings > Platform Integrations** 2. Click on **Hugging Face** 3. Enter your Hugging Face API token 4. Test the connection The platform creates environment variables automatically: - `HUGGINGFACE_TOKEN` ## Example Use Case: Prompt Injection Detector A Hugging Face guardrail model configured for prompt injection detection demonstrates the complete workflow: ### Arguments Configuration: - `text` (String, required) - `threshold` (Numerical, optional, default: "0.5") - `max_length` (Numerical, optional, default: "512") ### Usage: ```python # Model becomes available as: prompt_injection_detector result = prompt_injection_detector( text="Ignore previous instructions", threshold=0.8, max_length=512 ) ``` ## Want to Learn More? - Browse models at [Hugging Face Hub](https://huggingface.co/) - Review Hugging Face Transformers documentation - Check model licensing for compliance - Monitor usage through platform analytics --- # OpenAI Integration Source: https://docs.genguardx.ai/integrations/llm-providers/openai/ Markdown: https://docs.genguardx.ai/integrations/llm-providers/openai/index.md Description: Configure OpenAI in GGX and register OpenAI-backed models with API credentials, model settings, arguments, and governed scoring logic. The OpenAI integration provides direct access to OpenAI models through a unified interface. Configure once, use everywhere with enterprise-grade performance and the latest AI capabilities from OpenAI. ## Integrating OpenAI Simply enter your OpenAI API key once in the Platform Integrations section. This enables authorized users to access OpenAI models within the platform. Once integrated, models can be registered and used as any other python object on the platform. ```python # Example: Using a registered OpenAI model result = openai_gpt4_model(text="Analyze this data", temperature=0.8) ``` ## Example of Models Supported OpenAI provides access to state-of-the-art language models with different capabilities: **GPT-4o** - Most advanced multimodal model with vision and reasoning capabilities **GPT-4** - Advanced reasoning and complex task completion **GPT-3.5 Turbo** - Fast, efficient responses for most use cases **o1-preview, o1-mini** - Latest reasoning models with enhanced problem-solving **Additional Models** - Latest OpenAI variants with enhanced capabilities ## Registering a New OpenAI Model Navigate to **New Model** to begin registration. The registration form connects your OpenAI integration with custom model configurations. ### Basic Information **Description**: Document your model's purpose, use cases, and limitations. For example: "GPT-4o optimized for content analysis and generation. Use for enterprise content processing with multimodal capabilities. Ideal for complex reasoning and creative tasks." ### Code Configuration **Alias**: A unique identifier for your model (e.g., `openai_gpt4_analyzer`, `content_generator`). This becomes the variable name you'll use in code. **Output Type**: Define the return format: - `Map[String, String]` - Key-value pairs for structured responses - `String` - Simple text responses - `List` - Array of items **Input Type**: Select your implementation approach: - **API Based**: Platform handles API calls automatically using your OpenAI integration - **Python Function**: Custom function implementation with full control - **Custom**: Advanced configurations for specialized use cases **Model Provider**: Select "OpenAI" from your configured integrations. ### Arguments Configuration Define input parameters that your model will accept. Important: Variables declared here are automatically available in the Scoring Logic section. Common argument patterns for OpenAI models: | Alias | Type | Optional | Default Value | Usage | |-------|------|----------|---------------|-------| | `text` | String | No | N/A | Main input content | | `temperature` | Numerical | Yes | 0.7 | Controls response creativity | | `max_tokens` | Numerical | Yes | 1500 | Maximum response length | | `system_prompt` | String | Yes | "" | System instructions | Use **+ Add Argument** to include additional parameters. ### Scoring Logic Implementation In the Scoring Logic section, you can directly reference any variable declared in the Arguments section. The platform automatically makes these available in your code. ```python # Arguments: text, temperature are automatically available import os from openai import OpenAI # Direct initialization client = OpenAI( api_key=os.getenv("OPENAI_API_KEY") ) if text is None: return None messages = [ {"role": "system", "content": system_prompt if system_prompt else "You are a helpful assistant."}, {"role": "user", "content": text} ] completion = client.chat.completions.create( model="gpt-4o", messages=messages, max_tokens=int(max_tokens), temperature=float(temperature), top_p=0.95, frequency_penalty=0, presence_penalty=0 ) return {"output": completion.choices[0].message.content, "context": None} ``` ## Platform Integration Setup Before registering models, configure your OpenAI credentials: 1. Navigate to **Settings > Platform Integrations** 2. Click on **OpenAI** 3. Enter your OpenAI API key 4. Test the connection The platform creates environment variables automatically: - `OPENAI_API_KEY` ## Example Use Case: Content Analysis Model An OpenAI GPT-4o model configured for enterprise content analysis demonstrates the complete workflow: ### Arguments Configuration: - `text` (String, required) - `temperature` (Numerical, optional, default: "0.3") - `max_tokens` (Numerical, optional, default: "2000") - `system_prompt` (String, optional, default: "You are an expert content analyst.") ### Usage: ```python # Model becomes available as: content_analyzer result = content_analyzer( text="Your content text here...", temperature=0.3, max_tokens=2000, system_prompt="Analyze this content for key themes, sentiment, and actionable insights." ) ``` ## Want to Learn More? - Review [OpenAI API documentation](https://platform.openai.com/docs) - Check [OpenAI model capabilities](https://platform.openai.com/docs/models) - Monitor usage through OpenAI Dashboard - Set up rate limiting and cost management --- # Observability Source: https://docs.genguardx.ai/integrations/observability/ Markdown: https://docs.genguardx.ai/integrations/observability/index.md Description: Connect AI observability traces from LangSmith, Arize Phoenix, Langfuse, Humanloop, Datadog, and other tools to GGX for judges, human review, ground truth, bug categorization, and closed-loop AI lifecycle governance. AI observability tools help teams collect traces, inspect runs, measure latency and cost, evaluate outputs, and understand production behavior. GGX can sit alongside these tools by connecting their traces into the GGX monitoring and review workflow. The benefit is not just another dashboard. GGX helps make the AI lifecycle seamless: production traces become review queues, review outcomes become ground truth, common failures become categorized bugs, and the same evidence feeds future judges, evaluations, approvals, and deployment decisions. ## Supported observability patterns | Tool | Typical use | GGX connection pattern | | --- | --- | --- | | [LangSmith](langsmith/) | Trace, debug, evaluate, and monitor LangChain and broader LLM applications. | Export or stream traces into GGX monitoring, then run judges and human reviews on selected interactions. | | [Arize Phoenix](arize-phoenix/) | OpenTelemetry-based AI tracing, evaluations, prompt iteration, datasets, and experiments. | Bring traces, spans, evaluator outputs, and human labels into GGX as monitoring evidence and reusable test cases. | | [Langfuse](langfuse/) | Open-source LLM observability, prompt management, evaluations, dashboards, and annotation queues. | Connect traces and scores to GGX so monitoring outcomes can become ground truth, bugs, and approval evidence. | | [Humanloop](humanloop/) | Evaluation, prompt management, observability, human feedback, and production log review. | Import logs and review outcomes into GGX for governed SME review, judge creation, and lifecycle evidence. | | [Datadog LLM Observability](datadog/) | Monitor LLM applications alongside broader infrastructure, APM, logs, and security telemetry. | Connect LLM traces or incidents to GGX so AI-specific quality review feeds testing, approvals, and remediation workflows. | ## What GGX adds Observability platforms can show what happened. GGX turns that signal into governed lifecycle action. 1. **Connect traces:** Import production traces, spans, prompts, responses, metadata, scores, and user feedback from the observability tool. 2. **Reduce the review funnel:** Use heuristics, automated checks, and LLM judges to identify which traces need human attention. 3. **Run human reviews:** Route selected traces to business SMEs, risk teams, model owners, or annotation queues. 4. **Promote positive reviews:** If a reviewed trace is acceptable, add it to ground truth so future tests and monitoring become more representative. 5. **Capture negative reviews:** If a reviewed trace fails, create a bug or finding with the relevant prompt, model, RAG context, trace metadata, and reviewer notes. 6. **Categorize common bugs:** Group repeated issues into failure categories such as hallucination, retrieval miss, bad refusal, toxicity, policy breach, tool error, latency, or user-intent mismatch. 7. **Improve judges:** Use reviewed traces and failure categories to build, calibrate, and validate LLM-as-a-judge reports. 8. **Close the loop:** Feed the resulting ground truth, bugs, judges, and reports back into refinement, approval, deployment, and ongoing monitoring. This turns monitoring from a static set of metrics into a practical operating system for managing and improving AI applications. ## Choosing what to send to GGX Send enough trace context for review and remediation: - Prompt, response, model, model version, route, and provider. - Retrieval context, citations, tool calls, agent steps, and error details. - User feedback, thumbs up/down, ratings, comments, escalation flags, and session metadata. - Existing observability scores such as toxicity, hallucination, latency, cost, groundedness, retrieval quality, or custom evaluator results. - Business context such as product area, customer segment, channel, workflow, policy, or approved application version. GGX can then combine this with its own registered assets, reports, approval workflows, and monitoring logic. ## Related pages - [LangSmith](langsmith/) - [Arize Phoenix](arize-phoenix/) - [Langfuse](langfuse/) - [Humanloop](humanloop/) - [Datadog](datadog/) --- # Arize Phoenix Source: https://docs.genguardx.ai/integrations/observability/arize-phoenix/ Markdown: https://docs.genguardx.ai/integrations/observability/arize-phoenix/index.md Description: Connect Arize Phoenix traces, evaluations, datasets, and human annotations to GGX for governed monitoring, review, bug categorization, and AI lifecycle improvement. [Arize Phoenix](https://arize.com/docs/phoenix) is an AI observability and evaluation platform with tracing, evaluations, prompt engineering, datasets, experiments, and human annotations. Phoenix supports OpenTelemetry-based tracing and can capture model calls, retrieval, tool use, and custom logic. ## How Phoenix fits with GGX Phoenix helps teams debug and improve AI applications with traces, evaluators, datasets, and experiments. GGX can use Phoenix trace data as production monitoring input and connect it to review, approval, and lifecycle governance. Use this pattern when: - Your application is already instrumented with Phoenix or OpenInference. - You want Phoenix traces and spans to become GGX monitoring cases. - You want evaluator scores and human annotations to feed GGX ground truth. - You want recurring Phoenix failure patterns to become categorized GGX bugs or findings. ## Trace-to-review workflow 1. Send Phoenix traces, spans, evaluator scores, dataset references, and human labels into GGX. 2. Use GGX rules and LLM judges to filter the trace volume into a smaller review queue. 3. Ask business SMEs or risk reviewers to confirm whether selected interactions are acceptable. 4. Add positive reviews to ground truth for future simulations and judge calibration. 5. Add negative reviews to GGX bugs or findings with failure categories and linked trace context. 6. Use those categories to track recurring issues and prove that fixes reduce failure frequency. ## GGX value add Phoenix can show what happened inside an AI run and help teams evaluate it. GGX turns that evaluation signal into an operating process: reviewed traces become reusable ground truth, failures become managed remediation work, and monitoring becomes part of the same lifecycle used for refinement, approvals, deployment, and oversight. --- # Datadog Source: https://docs.genguardx.ai/integrations/observability/datadog/ Markdown: https://docs.genguardx.ai/integrations/observability/datadog/index.md Description: Connect Datadog LLM Observability traces and incidents to GGX for AI judges, human review, ground truth creation, bug categorization, and lifecycle governance. [Datadog LLM Observability](https://docs.datadoghq.com/llm_observability/) helps teams monitor LLM applications alongside infrastructure, APM, logs, security, dashboards, and alerting. It is useful when AI behavior needs to be understood together with production system health. ## How Datadog fits with GGX Datadog is often the operational system of record for production health. GGX can connect Datadog LLM traces, alerts, incidents, or linked logs into AI-specific monitoring and review workflows. Use this pattern when: - Your SRE or platform team monitors AI applications in Datadog. - LLM issues need to be reviewed by product, risk, compliance, or model owners. - Datadog alerts should trigger GGX review queues or annotation workflows. - Production traces should feed GGX ground truth, bugs, judges, and approval evidence. ## Trace-to-review workflow 1. Send Datadog LLM traces, alerts, incidents, logs, service metadata, latency, cost, model, prompt, response, and error context to GGX. 2. Use GGX judges and rules to separate normal telemetry from interactions that need human review. 3. Route the narrowed review set to the right SMEs, risk reviewers, or annotation queues. 4. Add accepted interactions to ground truth. 5. Add failed interactions to bugs or findings, grouped by recurring failure category. 6. Use the resulting evidence to improve prompts, retrieval, model routing, guardrails, and deployment readiness. ## GGX value add Datadog tells teams when something happened in production. GGX helps teams decide what that event means for AI quality, risk, approvals, and future releases. The result is a feedback loop where monitoring improves the AI system rather than only reporting metrics about it. --- # Humanloop Source: https://docs.genguardx.ai/integrations/observability/humanloop/ Markdown: https://docs.genguardx.ai/integrations/observability/humanloop/index.md Description: Connect Humanloop logs, evaluations, prompt management, and human feedback workflows to GGX for closed-loop AI lifecycle governance. [Humanloop](https://humanloop.com/docs) provides LLM evaluation, prompt management, observability, logs, human feedback, and reviewer workflows. Humanloop's documentation notes that the Humanloop platform will be sunset on September 8, 2025, so teams should confirm their current deployment and migration path before building new integrations. ## How Humanloop fits with GGX Humanloop logs and evaluations can provide useful review and product-feedback signals. GGX can ingest those signals and connect them to governed monitoring, ground truth, bug tracking, judge creation, and approval evidence. Use this pattern when: - Historical Humanloop logs contain useful production examples. - Humanloop human feedback or evaluator results should be preserved in GGX. - Your team is migrating review workflows into GGX. - You want production review outcomes to feed ongoing GGX monitoring and evaluation. ## Trace-to-review workflow 1. Export Humanloop logs, prompts, evaluation results, user feedback, reviewer labels, and metadata. 2. Import the records into GGX monitoring or datasets. 3. Run GGX judges to triage which records need human review. 4. Promote positive reviews to ground truth. 5. Convert negative reviews into bugs or findings with failure categories. 6. Use the resulting examples to create or improve GGX judges, reports, and simulations. ## GGX value add Humanloop-style feedback loops are strongest when they affect future development and deployment decisions. GGX keeps that loop active: monitoring evidence becomes test data, judge calibration, bug categories, approval evidence, and deployment guidance in one lifecycle. --- # Langfuse Source: https://docs.genguardx.ai/integrations/observability/langfuse/ Markdown: https://docs.genguardx.ai/integrations/observability/langfuse/index.md Description: Use Langfuse traces, scores, datasets, and annotation queues with GGX to close the monitoring loop across judges, human review, ground truth, bugs, and approvals. [Langfuse](https://langfuse.com/docs) is an open-source AI engineering platform for LLM observability, prompt management, evaluations, dashboards, datasets, experiments, user feedback, and annotation queues. ## How Langfuse fits with GGX Langfuse can capture production traces and scores across LLM calls, retrieval, embedding, API calls, sessions, agents, and user interactions. GGX can connect those traces to governed monitoring and review workflows. Use this pattern when: - You already use Langfuse as the trace and evaluation layer for LLM applications. - You want selected Langfuse traces to be reviewed by SMEs or risk teams in GGX. - You want Langfuse scores and user feedback to influence GGX monitoring rules. - You want reviewed Langfuse examples to power future GGX datasets, judges, and approval evidence. ## Trace-to-review workflow 1. Send Langfuse traces, scores, sessions, prompt versions, user feedback, and metadata to GGX. 2. Use GGX monitoring logic to identify traces that need judge review, human review, or no review. 3. Promote positively reviewed traces into ground truth. 4. Convert negatively reviewed traces into GGX bugs or findings. 5. Categorize common failures so teams can see where prompt, RAG, model, tool, or policy changes are needed. 6. Reuse those examples in GGX simulations and judge-development workflows. ## GGX value add Langfuse provides strong observability and evaluation workflows. GGX makes the monitoring lifecycle actionable across stakeholders: positive monitoring reviews become test data, negative reviews become remediation work, and both feed the evaluation and approval system used to manage production AI. --- # LangSmith Source: https://docs.genguardx.ai/integrations/observability/langsmith/ Markdown: https://docs.genguardx.ai/integrations/observability/langsmith/index.md Description: Use LangSmith traces with GGX to run judges, human reviews, ground truth promotion, bug tracking, and closed-loop AI monitoring. [LangSmith](https://docs.smith.langchain.com/) provides observability for LLM applications, including traces, production metrics, dashboards, alerts, automations, online evaluations, annotation, and user feedback. ## How LangSmith fits with GGX LangSmith is often the first place engineering teams inspect detailed traces for LangChain or other LLM applications. GGX can consume those traces as monitoring evidence and add lifecycle governance around them. Use this pattern when: - Engineering teams already instrument production applications with LangSmith. - You want business, risk, or compliance reviewers to evaluate selected traces in GGX. - You want positive production examples to become ground truth. - You want negative examples to become categorized bugs or findings. - You want LangSmith monitoring to feed future GGX judges, simulations, and approval evidence. ## Trace-to-review workflow 1. Export or stream LangSmith traces into GGX with prompt, response, model, metadata, evaluator scores, user feedback, and trace links. 2. Use GGX monitoring rules and judges to prioritize which traces need review. 3. Route selected traces to GGX human review or annotation queues. 4. Promote accepted traces to ground truth. 5. Convert rejected traces into bugs or findings, then categorize recurring failure patterns. 6. Use reviewed examples to improve GGX judges, reports, and regression test datasets. ## GGX value add LangSmith helps teams observe and debug LLM behavior. GGX closes the lifecycle loop: monitoring evidence becomes ground truth, bugs, judge calibration data, approval evidence, and deployment guidance. That makes production monitoring useful for the way teams manage AI systems, not just the way they inspect metrics. --- # LLM Judges Source: https://docs.genguardx.ai/llm-judges/ Markdown: https://docs.genguardx.ai/llm-judges/index.md Description: A gallery of ready-to-use LLM-as-a-Judge evaluators for GenGuardX. Browse by category and inspect each judge's prompt and code. ## What is an LLM Judge? An **LLM-as-a-Judge** uses a large language model to evaluate the output of another AI system — scoring qualities that are hard to measure with rules or exact-match metrics, such as answer relevancy, factual accuracy, coherence, toxicity, or bias. Instead of asking a human reviewer to grade every response, a judge is given a rubric (the **prompt**) and returns a structured verdict — a score and a short reasoning — that you can aggregate, monitor, and act on at scale. Every judge in this catalog is a ready-to-use GenGuardX component. It pairs a carefully written evaluation **prompt** with a `Judge Model` that is model-agnostic (it connects through **AnyLLM**, so you can point it at OpenAI, Gemini, Anthropic, Bedrock, and more) and validates its output against a **pydantic** model, so you always get well-formed, typed results. Search or filter by category below, then open any card to inspect its prompt and code. ## Syncing a judge to your instance Each judge ships as a self-contained Python file that declares its prompt and model with GGX decorators and syncs them with `ggx.sync`. To add a judge to your GenGuardX instance: 1. **Install the SDK** and open the judge you want from the catalog below. ```bash pip install genguardx ``` 2. **Copy the judge's code** into a file in your project (e.g. `judge.py`). Use the **Code** tab on any card to grab the full source. 3. **Set your credentials.** The judges read them from the environment — put your instance URL and [API key](/register-and-refine/sync/#obtaining-your-api-key) in a `.env` file: ```bash GGX_API_URL="https://your-ggx-instance.example.com" GGX_API_KEY="your-api-key-here" ``` 4. **Run the file to sync it.** The `__main__` block calls `ggx.init(...)` and `ggx.sync(...)`, registering the judge's prompt and model on your instance: ```bash python judge.py ``` Once synced, the judge appears in your inventory and can be used for downstream applications. See the full reference for details. GGX Sync reference :::caution[Test and adapt before you rely on a judge] These judges are strong, general-purpose starting points — not drop-in ground truth. LLM judges can be sensitive to your domain, data distribution, and scoring conventions. Before you depend on one, **validate it on your specific task and data** (ideally against a set of human-labeled examples) and **adapt the prompt, scoring scale, or model as needed** to get reliable results. ::: ## Browse the catalog --- # GenerativeAI Lifecycle Management Source: https://docs.genguardx.ai/register-and-refine/ Markdown: https://docs.genguardx.ai/register-and-refine/index.md Description: Use GGX to manage the GenAI lifecycle from use case definition and data collection through pipeline development, evaluation, approval, deployment, and production monitoring. ## Generative AI Lifecycle: The Generative AI development lifecycle is a systematic framework designed to help businesses create effective AI-driven solutions across various applications. Each phase is vital in ensuring the technology’s accuracy, dependability, and ethical implementation. ![GenAI Lifecycle](./genai-lifecycle.excalidraw.svg) Below is a breakdown of each stage and its significance: #### **1. Business Use Case Identification** The first step in the lifecycle involves defining the problem statement and objectives that the solution aims to address. Clearly defining the business goals is a very crucial first step as it lays the foundation for a well-structured development and validation strategy. #### **2. Training and Validation Data Collection** A GenAI pipeline relies heavily on high-quality, diverse datasets that cover all the business scenarios for testing and validation. Data must be gathered from reliable sources, preprocessed for consistency, and cleaned to remove noise. The platform enforces strong data governance and data quality checks that ensure integrity, compliance, and fairness, which directly impact model performance and validation. Usually, sample data is created out of this for quick testing during the pipeline development phase to expedite iterations. #### **3. Pipeline Development** Developing an AI pipeline is an iterative process aimed at rapidly building the first version of prompts, models, and Retrieval-Augmented Generation (RAG) components to address a specific business use case. Once the initial version is ready, incremental improvements can be made by refining each component and analyzing performance gains. - **Crafting the Initial Prompt** Designing an effective prompt is crucial for generating accurate and relevant responses. However, prompt engineering is an iterative process that requires multiple refinements as development progresses. - **Connecting to External Knowledge** Enhancing model responses with external knowledge sources improves contextual accuracy. In GenAI pipelines, this is achieved using RAG techniques. Building a robust RAG architecture ensures the pipeline retrieves the most relevant information, significantly impacting overall accuracy. - **Choosing the Best LLM for the Use Case** Selecting an optimal LLM depends on various factors such as the business problem, expected outcomes, cost, and operational efficiency. Choosing the right model ensures the pipeline aligns with business objectives. - **Pipeline Creation** Once the foundational components are established, end-to-end pipeline development begins. Often creating multiple pipeline versions and comparing them helps in selecting the best-performing one. One can swap in and swap out multiple building blocks of the pipeline to get to the best state. - **Experimentation with Algorithms and Settings** Optimizing a pipeline requires experimenting with different models, algorithms, and hyperparameter settings (e.g., temperature, top-k, top-p). Fine-tuning these parameters ensures the model generates responses that align with performance goals. #### **4. Evaluations - Automated and Human-based** Before deploying a GenAI pipeline, it must be rigorously tested for accuracy, fairness, stability and robustness. A combination of automated evaluation techniques and human assessments ensures comprehensive validation. Tracking performance metrics and refining the model based on evaluation results helps maintain reliability and transparency. #### **5. Move to Production** Once the pipeline is validated, it is deployed into production environments where it integrates with existing business systems so that it can be used for real-world decision-making. #### **6. Production Monitoring** Continuous monitoring is crucial to maintaining performance and mitigating issues such as response drift and biases. Implementing monitoring tools ensures sustained accuracy and compliance with business and regulatory requirements. ## How does GGX help with GenAI Lifecycle Management? - **Data Integration:** Supports validation and testing data registration by integrating with production data lakes and in-house storage solutions. - **Component Inventory:** Maintains a registry of essential GenAI pipeline components, including models, RAGs, prompts, and use-case-specific pipelines to enhance reusability and develop smaller building blocks. - **Governance and Collaboration:** Provides features such as version management, change history, approvals, ongoing reviews, lineage tracking, impact assessment, and team collaboration to enhance governance, experiment tracking, MRM, FL and business approvals. - **Evaluation and Reporting:** Enables assessment of complete pipelines and individual components using standardized (tailored for MRM, Fair Lending, and business needs) and custom reports. - **Human-Integrated Pre-Production Testing:** Facilitates real-world-like testing by allowing different teams to conduct robust validation before deployment. - **Artifact Export and Documentation:** Supports production artifact export and ODD generations. - **Monitoring:** Connects to production data sources for response labelling and automated performance monitoring and alerting. --- # Collaboration Source: https://docs.genguardx.ai/register-and-refine/collaboration/ Markdown: https://docs.genguardx.ai/register-and-refine/collaboration/index.md Description: How teams collaborate in GGX — sharing and requesting object access, role-based access management, external sync, groups, workspaces, and monitoring dashboards. GGX is built to help teams **develop, test, and monitor** GenAI pipelines together. Assets from different workstreams live in one central place, so developers, reviewers, and testers can build, iterate, and reuse each other's work instead of starting from scratch. ## Sharing and requesting access Any registered object in **Draft** or **Pending Approval** status can be shared with other users — or a user can raise a request for it. Access comes in two levels: View the object, run evaluations, download its documentation and artifacts, and **reuse** it inside another object. Everything Read allows, plus **edit, delete, send for approval, and share** the object with others. When requesting access, you choose the **scope**: the object on its own, or the object **together with its lineage** — every upstream component it depends on. :::note[Re-share when lineage changes] If you add new objects to something already shared with full lineage, you must **re-share** it so collaborators keep access to the complete dependency chain. ::: ### Rules worth knowing | Rule | Detail | |------|--------| | **Eligibility** | You can only request access to objects your **user role** (set in Settings) makes you eligible for. | | **Read can't re-share** | A non-owner with **Read** only cannot share the object onward. | | **Notifications** | When an object is shared, the receiver gets a notification that it is now available to them. | | **Revocable** | Shared access can be changed or revoked at any time by the **owner** or anyone with **Write** access. | | **Notebook sharing** | Access can also be shared from a **Jupyter Notebook** via the GGX package. | ## Access management **Role-Based Access Management** governs what each user can do. Every onboarded user is assigned a **role**, configured under **Settings → Roles**. A role is built from three nested controls — you set how much of the platform a user sees, then exactly what they can do with the objects inside it.
![Role-based access as three nested controls: Module Access (Enable, Disable, Hidden) contains Pages Access (Enable, Disable, Hidden), which contains Object-Level Authority (Read, Write, Approve) — elevated with Superuser, scoped to a Collection, or refined with Additional Specifications rules.](./access-scope.svg)
A role nests from whole-module visibility down to per-object authority on a single collection.
| Control | What you set | Options | |---------|--------------|---------| | **Module Access** | Visibility of each module — Data & AI Assets, Human Integrated Testing, Risk & Compliance, Monitor & Track, Settings. | **Enable** · **Disable** · **Hidden** | | **Specific Pages Access** | Visibility of each page within an enabled module — Table Registry, Projects, Model Catalog, Prompt Registry, RAG Registry, Pipeline Registry, Global Functions, Reports. | **Enable** · **Disable** · **Hidden** | | **Object-Level Authority** | What the role can do with each object type — Data Table, Project, Model, Prompt, and so on. | **Read** · **Write** · **Approve** | Object-Level Authority can then be **elevated** and **scoped** per object type: - **Superuser** — elevate any authority to full control of that object type. - **Collection** — restrict the authority to a named **collection** of objects rather than all of them. - **Additional Specifications** — refine access with custom rules built from **LHS → Operator → RHS** conditions on an object's properties (for example, only objects whose Group equals a given value). :::note[Allow Data Upload] A role can also carry standalone permissions such as **Allow Data Upload**, which lets its users supply custom data in simulations. ::: ## Integrating external updates A registered object's definition can be **exported**, modified outside the platform, and **re-synced** using GGX commands. GGX automatically tracks and records every external change, keeping the history clear and consistent. ## Organizing work across teams Three mechanisms keep many teams working in the same platform without stepping on each other: Classify objects for **control and display**. Groups are object-type-specific: teams create custom groups visible only to their members, and registries display objects by group so they're easy to find. Administrators can also use groups to assign access when defining roles. Create **multiple independent workspaces** so teams can work in isolation, without visibility into each other's in-progress work. ## Monitoring and alerting Build **customized dashboards** for key stakeholders and leadership — a bird's-eye view of activity across teams and across every stage of the application lifecycle, with alerting on the metrics that matter. --- # Pipeline Registration Guide Source: https://docs.genguardx.ai/register-and-refine/examples/intent-classification-pipeline-registration/pipeline/ Markdown: https://docs.genguardx.ai/register-and-refine/examples/intent-classification-pipeline-registration/pipeline/index.md Description: Register an intent classification pipeline in GGX by connecting a model and prompt, configuring pipeline metadata, adding resources, and testing outputs. A pipeline combines multiple resources (models, prompts, RAGs, helper functions) to create an end-to-end use-case specific workflow. Read more about [Pipelines](../../../inventory-management/pipelines/) to understand more about what they are and how they work. This guide covers how to register pipelines on the GGX, using an **Intent Classification Pipeline** as a working example. ## Prerequisites Before registering a pipeline, ensure you have: - ✅ **Registered a Model** - Follow the [Model Registration Guide](../../model/) to register Gemini 2.0 Flash - ✅ **Registered a Prompt** - Follow the [Prompt Registration Guide](../prompt/) to register the intent classification prompt **Quick Check:** Navigate to **GenAI Studio → Model Catalog** and **Prompt Registry** to verify your resources are available. If you haven't completed these steps, please do so before proceeding. --- ## Registration Steps ### Step 1. Navigate to Pipeline Registry Go to **GenAI Studio → Pipeline Registry** and click the **Create** button. ### Step 2. Fill in Basic Information ![alt text](intent-pipeline-description.png) **Basic Information** fields help organize and identify your pipeline: - **Description:** Clear explanation of what the pipeline does and its workflow - **Usecase Type:** The primary use case category (e.g., "Question Answering") - **Task Type:** Specific task the pipeline performs (e.g., "Classification") - **Impact of Generated Output:** Scope of the pipeline's usage (e.g., "Internal Only") - **Data Usage:** Whether the pipeline uses additional data sources beyond user input - **Group:** Category for organizing similar pipelines (e.g., "Conversational AI ChatBot Pipeline") - **Permissible Purpose:** Approved use cases and business scenarios for this pipeline **Example for Intent Classification Pipeline:** ``` This is a chat-based pipeline designed for intent recognition & classifies incoming user messages by assigning them to one of the following predefined intents: 1. ACTIVATE CARD 2. APPLY FOR LOAN 3. APPLY FOR MORTGAGE 4. BLOCK CARD 5. CANCEL LOAN This pipeline operates on a single input and produces a single output for each message, without maintaining conversational context across interactions. ``` ### Step 3. Configure Code Settings ![alt text](intent-pipeline-code-configure.png) **Code Settings** define how your pipeline operates and which resources it uses. Fill in the **Basic Information** fields as shown in the image above. These includes Alias, Input Type, Pipeline Type, and Context Type. NOTE: For this case we have chosen to create a chat-based pipeline as in future we can expand the pipeline to recognize the user's intent over multiple turns, though for now we are keeping it simple and using a single input/output. ### Step 4. Add Resources **Resources** are the pre-registered components your pipeline will use. Click **+ Create New** or search for existing resources to add: **LLMs / Models:** - `gemini_2_0_flash` - The foundation model for generating responses **Prompts:** - `customer_intent_classification_prompt` - The structured instructions for intent classification **Other Resources** (Not required for this example): - **RAGs:** For retrieving relevant documents - **Agents & Sub-Pipelines:** For complex multi-step workflows - **Helper Functions:** For data processing utilities, or any other function according to your requirement ### Step 5. Write Pipeline Scoring Logic ![alt text](intent-pipeline-scoring.png) **Pipeline Scoring Logic** orchestrates how resources work together: - Combines models, prompts, and other resources - Processes user inputs and conversation history - Generates outputs and maintains context across turns **Example - Intent Classification Pipeline:** ```python import json # Run Gemini with prompt and user input response = gemini_2_0_flash( customer_intent_classification_prompt(user_message=user_message) ) # Parse the JSON response to get the classified intent response_json = json.loads(response["response"]) classified_intent = response_json.get("classified_intent", "UNKNOWN") # List of valid intents valid_intents = [ "ACTIVATE CARD", "BLOCK CARD", "CARD DETAILS", "CHECK CARD ANNUAL FEE", "CHECK CURRENT BALANCE ON CARD", ] # Validate the classified intent if classified_intent not in valid_intents: classified_intent = "UNKNOWN" # Return output and context return { "output": classified_intent, "context": "Any information that needs to be stored across turns" } ``` **What This Does:** 1. Calls the Gemini model with the classification prompt and user message. 2. Parses the JSON response to extract the classified intent. 3. Validates the intent against the list of valid intents. 4. Returns the classification result as the output and any information that needs to be stored across turns as the context. **Variables Available:** - `user_message` - The current user input (type: String) - `history` - Previous conversation messages (type: list[TypedDict[{'role': str, 'content': str}]]) - `context` - Any information that needs to be stored across turns (type: String) ### Step 6. Add Examples (Optional) ![alt text](intent-pipeline-examples.png) Add test examples to validate pipeline behavior. **Note:** Examples help with testing and documenting expected behavior. ### Step 7. Save the Pipeline Click **Create** to register the pipeline. The pipeline is now: - Available in the Pipeline Registry - Ready for simulation and testing - Ready for use in downstream applications --- ## Testing Your Pipeline After creating the pipeline, test it to verify behavior: ### Quick Test (During Creation/Editing) 1. While creating or editing the pipeline, scroll to the **Code** section 2. Click **Test Code** in the bottom right corner 3. Enter test inputs to verify logic without saving ### Interactive Test (After Saving) 1. Navigate to your saved pipeline 2. Click **Run** → **Chat Session** (top right corner) NOTE: Chat sessions is only available for chat-based pipelines. For free-flow pipelines, you can test the pipeline by calling the pipeline function with sample inputs using the test code button. 3. Enter sample messages to test the full conversation flow 4. Verify outputs match expected behavior --- ## Using Pipelines ### In Applications Reference the pipeline in your application code: ```python # Call the pipeline result = customer_intent_classification_pipeline( user_message="I want to block my card", history=[], context="" ) # Access the output classified_intent = result["output"] # "BLOCK CARD" context = result["context"] # "Any information that needs to be stored across turns" ``` --- ## Related Documentation - [Model Registration Guide](../../model/) - Register foundation models - [Prompt Registration Guide](../prompt/) - Create reusable prompts --- By following this guide, you can create reliable, production-ready pipelines that combine multiple AI resources into cohesive workflows. --- # Prompt Registration Guide Source: https://docs.genguardx.ai/register-and-refine/examples/intent-classification-pipeline-registration/prompt/ Markdown: https://docs.genguardx.ai/register-and-refine/examples/intent-classification-pipeline-registration/prompt/index.md Description: Register an intent classification prompt in GGX by defining prompt metadata, template variables, structured instructions, examples, and test cases. This guide covers how to register prompts on the GGX, using an **Intent Classification Prompt** as a working example. If you are new to Prompts, then this doc might help you understanding what they are and how do they work -> [Prompts](../../../inventory-management/prompts/) --- ## Registration Steps ### Step 1. Navigate to Prompt Registry Go to **GenAI Studio → Prompt Registry** and click the **Create** button. ### Step 2. Fill in Basic Information **Example for Intent Classification:** ![alt text](prompt-description.png) **Basic Information** fields help organize and identify your prompt: - **Description:** Clear explanation of what the prompt does and its purpose - **Group:** Category for organizing similar prompts (e.g., "Existing Customer Credit Card Related Prompts") - **Permissible Purpose:** Approved use cases and business scenarios for this prompt - **Task Type:** Classification of the prompt's function (e.g., "Classification" for intent detection) - **Prompt Type:** Format of the prompt (e.g., "System Instruction" for system-level prompts) - **Prompt Elements:** Optional tags or metadata for additional categorization ### Step 3. Configure Prompt Template ![alt text](prompt-template.png) **Alias:** `customer_intent_classification_prompt` - A Python variable name to reference this prompt in pipelines #### Example Prompt Template The **Prompt Template** is where you write the actual instructions for the LLM: - Use `{}` placeholders for dynamic variables (e.g., `{customer_utterance}`) - Write clear, structured instructions for the model to follow - Include examples to guide the model's behavior - Define expected output format (e.g., JSON schema) **Example Prompt Template for Intent Classification:** ````markdown # PERSONA & TONE You are a trusted, efficient, and security-conscious digital assistant, specialized in handling banking-related queries for existing customers of BankX. Maintain a tone that is: - Professional: Clear, formal, and polite - Concise: Direct answers without filler - Data-driven: Never guess; respond only based on verified data - English only # GOAL Accurately predict customer intent from a predefined list of possible intents. # TASK INSTRUCTIONS: ### Step 1: Review Intent Definitions Thoroughly understand the predefined list of intents. ### Step 2: Pre-Defined List of Intents #### ACTIVATE CARD - Definition: Request to activate a newly issued card - Examples: • "How do I activate my new debit card?" • "Activate my credit card now." #### BLOCK CARD - Definition: Request to block lost, stolen, or compromised card - Examples: • "Block my credit card immediately." • "I lost my debit card, can you block it?" #### CARD DETAILS - Definition: Inquiry about card information - Examples: • "How many cards do I have?" • "What is the name on my card?" #### CHECK CARD ANNUAL FEE - Definition: Inquiry about annual fees - Examples: • "What's the annual fee for my credit card?" • "How much is my card's yearly charge?" #### CHECK CURRENT BALANCE ON CARD - Definition: Inquiry about available balance - Examples: • "What's my credit card balance?" • "How much money is on my debit card?" ### Step 3: Disambiguate and Summarize Customer Utterance - Overlook grammatical/spelling errors - Ignore PII (name, age, gender, personal data) - Focus on main intention in long sentences ### Step 4: Mapping Query to Intent - Map to most suitable intent from predefined list - Ensure only one intent is chosen - Recheck classification is in predefined list ### Step 5: Schema Compliance OUTPUT FORMAT: ```json {{"classified_intent": "str"}} ``` # EXAMPLE SCENARIOS: Example 1: Input: "I need to activate my new credit card." REASONING STEPS: - Review intent definitions - Understand all available intents - No disambiguation needed (clear query) - Maps to "ACTIVATE CARD" intent - Output in JSON format Output: ```json {{"classified_intent": "ACTIVATE CARD"}} ``` # Customer Query Query: {customer_utterance} ```` #### Define Arguments Arguments are inputs that get passed into the prompt template. Click **+ Add Argument** to add: | Alias | Type | Is Optional | Default Value | | -------------- | ------ | ----------- | ------------- | | `user_message` | String | ☐ No | - | **Note:** Use `{customer_utterance}` in the template and map it from `user_message` in Prompt Creation Logic. ### Step 4. Write Prompt Creation Logic **Prompt Creation Logic** allows you to programmatically process arguments before they're inserted into the template. This is useful for: - Formatting complex data structures - Generating dynamic content (like the intent list) - Applying conditional logic based on inputs - Validating or transforming user inputs **Example - Formatting Intent Definitions:** ![alt text](prompt-creation.png) ```python intent_definitions = [ { "Intent": "ACTIVATE CARD", "Definition": "Request to activate a newly issued card", "Examples": [ "How do I activate my new debit card?", "Activate my credit card now.", ], }, { "Intent": "BLOCK CARD", "Definition": "Request to block a lost, stolen, or compromised card", "Examples": [ "Block my credit card immediately.", "I lost my debit card, can you block it?", ], }, { "Intent": "CARD DETAILS", "Definition": "Inquiry about card information", "Examples": [ "How many cards do I have?", "What is the name on my card?", ], }, { "Intent": "CHECK CARD ANNUAL FEE", "Definition": "Inquiry about annual fees", "Examples": [ "What's the annual fee for my credit card?", "How much is my card's yearly charge?", ], }, { "Intent": "CHECK CURRENT BALANCE ON CARD", "Definition": "Inquiry about available balance", "Examples": [ "What's my credit card balance?", "How much money is on my debit card?", ], }, ] def get_intent_info(data_list): """Format intent definitions into readable text""" formatted_list = [] intent_number = 1 for item in data_list: formatted_list.append(f"#### {intent_number}. {item['Intent'].upper()}") formatted_list.append(f"- Definition: {item['Definition']}") formatted_list.append(f"- Examples:") for example in item["Examples"]: formatted_list.append(f" • {example}") formatted_list.append("") # Empty line between intents intent_number += 1 return "\n".join(formatted_list) # Fill in the prompt template return prompt.format( customer_utterance=user_message, list_of_intents=get_intent_info(intent_definitions) ) ``` **What This Does:** 1. Defines 5 card-related intent definitions with examples 2. Formats them into a structured, numbered list 3. Fills in `{customer_utterance}` and `{list_of_intents}` placeholders ### Step 5. Save the Prompt Click **Create** to register the prompt. The prompt is now: - Available in the Prompt Registry - Usable in pipelines and other objects ### Analyze and Improve the Prompt using GGX Capability After saving the prompt, you can test and refine it directly within **GenAI Studio**: - **🔍 Analyze Prompt:** Click the **Analyze Prompt** button to evaluate how your prompt behaves with different inputs. This helps you confirm that argument mappings, placeholders, and output formats are working correctly. - **✨ Improve with AI:** Use the **Improve with AI** button to automatically optimize your prompt. This provides AI-generated suggestions to enhance clarity, tone, and structure — helping improve prompt performance and consistency. --- ## Using Prompts in Pipelines Once registered, prompts can be used in downstream applications: ```python # Reference the prompt in pipeline code intent_result = customer_intent_classification_prompt( user_message=user_input ) # Access the classified intent classified_intent = intent_result["classified_intent"] # Use in downstream logic if classified_intent == "ACTIVATE CARD": # Handle card activation pass elif classified_intent == "BLOCK CARD": # Handle card blocking pass ``` --- ## Next Steps After registering your prompt: 1. **Register a model** - If you haven't already, register the LLM to use with this prompt 2. **Build a pipeline** - Combine your prompt with a model and other resources to create a use-case specific pipeline. --- ## Related Documentation - [Model Registration Guide](../../model/) - Register LLM models to use with prompts --- --- # Pipeline Registration Guide: English to French Translation Source: https://docs.genguardx.ai/register-and-refine/examples/language-translation-pipeline-registration/pipeline/ Markdown: https://docs.genguardx.ai/register-and-refine/examples/language-translation-pipeline-registration/pipeline/index.md Description: Register an English-to-French translation pipeline in GGX using Gemini 2.0 Flash, custom translation logic, pipeline metadata, and usage tracking. This guide walks you through registering an **English to French Translation Pipeline** on the GGX. This pipeline automatically detects English text and provides high-quality French translations using Gemini 2.0 Flash. **What This Pipeline Does:** - Detects if input text is in English - Translates English text to French with preserved tone and style - Returns error messages for non-English input - Tracks API usage costs If you are new to Pipelines, read [What are Pipelines?](../../../inventory-management/pipelines/) to understand how they work. ## Prerequisites Before registering this pipeline, ensure you have: - ✅ **Registered Gemini 2.0 Flash Model** - Follow the [Model Registration Guide](../../model/) to register the model - ✅ **API Token Configured** - Ensure `GOOGLE_API_TOKEN` is set up in Platform Integrations **Quick Check:** Navigate to **GenAI Studio → Model Catalog** and verify `gemini_2_0_flash` is available. If you haven't completed these steps, please do so before proceeding. ## Registration Steps ### Step 1. Navigate to Pipeline Registry Go to **GenAI Studio → Pipeline Registry** and click the **Create** button. ### Step 2. Fill in Basic Information ![Pipeline Basic Information](english-to-french-pipeline-description.png) **Basic Information** fields help organize and identify your pipeline: - **Description:** Clear explanation of what the pipeline does and its workflow - **Usecase Type:** The primary use case category - select **Translation** - **Task Type:** Specific task the pipeline performs - select **Generative Responses** - **Impact of Generated Output:** Scope of the pipeline's usage - select **External Facing** - **Data Usage:** Whether the pipeline uses additional data sources - leave empty for this pipeline - **Group:** Category for organizing similar pipelines - select **Example Pipelines** - **Permissible Purpose:** Approved use cases and business scenarios for this pipeline **Example Description:** ``` English to French Translation Assistant - Translate English text to French using Gemini 2.0 Flash. Key Features: - Automatic English language detection - High-quality French translations using Gemini 2.0 Flash - Preserves tone, style, and cultural nuances - Cost tracking for API usage Usage: - Simple translation: "Hello, how are you?" - Any English text: "The weather is beautiful today." - Formal or informal: Automatically preserves the tone Note: This pipeline only translates FROM English TO French. ``` ### Step 3. Configure Code Settings ![Pipeline Code Configuration](english-to-french-pipeline-code-config.png) **Code Settings** define how your pipeline operates and which resources it uses. **Configuration Fields:** - **Alias:** `english_to_french_translation`: A Python variable name to reference this pipeline in code - **Input Type:** Select **Python Function** : This pipeline uses custom Python code for translation logic - **Agent Provider:** Select **Other** : We're not using a pre-built agent provider for this translation pipeline - **Pipeline Type:** Select **Chat Based Pipeline** :Enables conversational interface and message history. - **Context Type:** `dict[str, str]` - Data type for storing information across conversation turns - For this pipeline, context stores translation metadata (costs, language detection) - **Interaction Type:** `TypedDict[{'role': str, 'content': str}]` - Format for conversation history messages - Standard chat message format with role (user/assistant) and content 💡 *Note: While this is a single-turn translation, Chat Based Pipeline allows for future enhancements like multi-turn conversations* ### Step 4. Add Resources ![Pipeline Resources](english-to-french-pipeline-resource.png) **Resources** are the pre-registered components your pipeline will use. Click **+ Create New** or search for existing resources to add: **LLMs / Models:** `gemini_2_0_flash` - The foundation model for generating translations **Prompts (Optional):** `english_to_french_translation` - The translation prompt for generating translations - Follow the [Prompt Registration Guide](../../intent-classification-pipeline-registration/prompt/) to create a reusable prompt **Other Resources** (Not required for this pipeline): - **RAGs:** For retrieving translation dictionaries or context - **Agents & Sub-Pipelines:** For complex multi-step translation workflows - **Helper Functions:** For pre/post-processing text ### Step 5. Write Pipeline Scoring Logic ![Pipeline Scoring Logic](english-to-french-pipeline-scoring-logic.png) **Pipeline Scoring Logic** orchestrates how resources work together to perform the translation. **Variables Available in the Pipeline:** - `user_message` - The English text to translate (type: String) - `history` - Previous conversation messages (type: list[TypedDict[{'role': str, 'content': str}]]) - `context` - Information stored across turns (type: dict[str, str]) **Complete Pipeline Code:** ```python # Step 1: Generate strict translation prompt prompt = english_to_french_translation(user_message=user_message) # Step 2: Get translation from Gemini result = gemini_2_0_flash( text=prompt, temperature=0.3, system_instruction='None' ) translated_text = result["response"] # Step 3: Return result return { "output": translated_text } ``` **What This Code Does:** - Use the registered Translation Prompt to convert the user's message to a French translation - Call the registered Gemini 2.0 Flash Model with the translation prompt to generate the translation: - Return the translation as the output of the pipeline ### Step 6. Add Examples (Optional) ![Pipeline Examples Section](english-to-french-pipeline-examples.png) Add test examples to validate pipeline behavior: | Input | Expected Output | |-------|----------------| | "Hello, how are you?" | "Bonjour, comment allez-vous ?" | | "The weather is beautiful today." | "Le temps est magnifique aujourd'hui." | | "Thank you very much!" | "Merci beaucoup !" | | "Hola, ¿cómo estás?" (Spanish) | "Error: Input text must be in English. Detected language: Spanish" | **Note:** Examples help with testing and documenting expected behavior. They also serve as regression tests when updating the pipeline. ### Step 7. Save the Pipeline Click **Create** to register the pipeline. The pipeline is now: - ✅ Available in the Pipeline Registry - ✅ Ready for simulation and testing - ✅ Ready for use in downstream applications ## Testing Your Pipeline After creating the pipeline, test it to verify translation quality and error handling ### Quick Test (During Creation/Editing) ![Test Code](english-to-french-pipeline-test-code.png) 1. While creating or editing the pipeline, scroll to the **Code** section 2. Click **Test Code** in the bottom right corner 3. Enter test inputs to verify logic without saving **Sample Test Cases:** ```python # Test Case 1: Simple greeting user_message = "Hello, how are you?" # Expected: "Bonjour, comment allez-vous ?" # Test Case 2: Non-English input (error handling) user_message = "Hola, ¿cómo estás?" # Expected: "Error: Input text must be in English. Detected language: Spanish" ``` ### Interactive Test (After Saving) - Navigate to your saved pipeline - Click **Run** → **Chat Session** (top right corner) - Enter sample English messages to test the translation flow **Verify:** - Translations are accurate and natural - Tone and style are preserved (formal/informal) - Non-English inputs return proper error messages - Output format is clean (no extra explanations) **Chat Session Testing Tips:** - Test both formal and informal language - Try technical terms and idioms - Verify cultural nuances are preserved - Test edge cases (very short/long text, special characters) ## Want to Improve/Extend Your Pipeline? Try These Ideas: - Auto-detect source language using NLP techiques and see how it performs compared to the current pipeline - Add translation confidence scores and quality of the translations using evaluation providers - Extend to support other languages ## Conclusion: You've successfully learned how to register an English to French Translation Pipeline that: - ✅ Detects English language automatically - ✅ Provides high-quality French translations - ✅ Handles non-English input gracefully - ✅ Maintains clean, production-ready code with reusable translation prompt and Gemini 2.0 Flash model ## Related Documentation - [Model Registration Guide](../../model/) - Register foundation models like Gemini 2.0 Flash - [Prompt Registration Guide](../../intent-classification-pipeline-registration/prompt/) - Create reusable prompts --- # Model Registration: Gemini 2.0 Flash Source: https://docs.genguardx.ai/register-and-refine/examples/model/ Markdown: https://docs.genguardx.ai/register-and-refine/examples/model/index.md Description: Register Gemini 2.0 Flash in the GGX Model Catalog with provider settings, input arguments, scoring logic, model metadata, and test examples. This guide covers registering the Gemini 2.0 Flash model on the platform. **Gemini 2.0 Flash** is Google's language model for classification and structured output tasks. --- ## Registration Steps ### Step 1. Navigate to Model Catalog Go to **GenAI Studio → Model Catalog** and click the **Create** button. ### Step 2. Fill in Basic Information ![alt text](model-description.png) **Basic Information** fields help organize and identify your model: - **Name:** Human-readable identifier for the model (e.g., "Gemini 2.0 Flash") - **Description:** Brief explanation of the model's purpose and capabilities - **Group:** Category for organizing similar models together (e.g., "Foundation LLMs") - **Permissible Purpose:** Approved use cases and business scenarios for this model - **Ownership Type:** License type - Proprietary, Open Source, or Internal - **Model Type:** Classification of the model (e.g., "LLM" for language models) ### Step 3. Configure Inferencing Logic #### Choose Input Type **Input Type:** You have two options: - **API Based** - Use this when working with models through API providers (OpenAI, Anthropic, Google Vertex AI, etc.) - **Python Function** - Use this for custom Python implementations or local models For this guide, we'll use **API Based**. #### Select Model Provider **Model Provider:** Select `Google Vertex AI` from the dropdown Once you select a provider, additional fields will appear to configure how the model is called: ![alt text](model-code-configure.png) - **Alias:** Variable name to reference this model in pipeline code (e.g., `gemini_2_0_flash`) - **Output Type:** Data type returned by the model (e.g., `dict[str, str]`) - **Input Type:** Choose between API-based (for external providers) or Python Function (for custom code) - **Model Provider:** Select the API provider hosting the model (Google Vertex AI) - **Model:** Specific model version from the provider's catalog (Gemini 2.0 Flash) #### Define Arguments The inputs to the model - messages, system instruction, temperature, etc. Click **+ Add Argument** to add each argument: | Alias | Type | Is Optional | Default Value | |-------|------|-------------|---------------| | `text` | String | ☐ | - | | `temperature` | Numerical | ☑ | 0 | | `system_instruction` | String | ☑ | None | **Argument Descriptions:** - `text`: The input prompt to send to the model - `temperature`: Controls randomness (0 = deterministic, 1 = creative) - `system_instruction`: Optional system-level instructions for the model You can add additional arguments based on your model's requirements. #### Write Scoring Logic ![alt text](model-scoring.png) Provide logic to initialize and score the model: ```python import os from google import genai from google.genai import types client = genai.Client(api_key=os.getenv("GOOGLE_API_TOKEN")) config = types.GenerateContentConfig( temperature=temperature, seed=2025, system_instruction=system_instruction ) response = client.models.generate_content( model="gemini-2.0-flash", contents=text, config=config ) return { "response": response.text, } ``` **What This Code Does:** - Authenticates using the `GOOGLE_API_TOKEN` environment variable (configured in Platform Integrations) - Sets up generation config with temperature and system instruction - Calls the Gemini 2.0 Flash model with the input text - Returns the generated response ### Step 4. Save the Model Add any notes or additional information in the **Additional Information** section, then click **Create** to complete registration. ### Step 5. Quick Example Run Click **Test Code** to run a sample query. ![alt text](model-test-code.png) Use the platform's test interface to verify: - Verify API authentication is working - Test with sample inputs before using in production - Debug any configuration issues - Validate the output format matches expectations ## Usage in Pipelines Once registered, the model appears in your Resources library and can be selected for any downstream usages. **Reference in pipeline code:** ```python # Call the registered model response = gemini_2_0_flash( text=user_prompt, temperature=0.7, system_instruction="You are a helpful assistant." ) # Access the response output_text = response["response"] ``` --- ## Related Documentation - [Prompt Registration Guide](../intent_classification_pipeline_registration/prompt/) - Create reusable prompts - [Google Gemini API Docs](https://ai.google.dev/gemini-api/docs) - Official Google documentation --- --- # Inventory Management Source: https://docs.genguardx.ai/register-and-refine/inventory-management/ Markdown: https://docs.genguardx.ai/register-and-refine/inventory-management/index.md Description: Register, organize, govern, share, version, and monitor GGX data and GenAI assets including tables, models, prompts, RAGs, and end-to-end pipelines. ## Overview The platform allows registering, tracking, and monitoring of Data and GenAI assets (like RAG, Models, LLMs and Pipelines) at a centralized location. ## Why Inventory Management is Helpful? - **Governance and Compliance:** Helps track AI models, datasets, and dependencies for regulatory audits. - **Reusability & Efficiency**: Prevents duplication of efforts by enabling teams to reuse registered and approved assets and standardized inventories reducing onboarding time for new teams. - **Security & Access Control**: Centralized inventories allow proper role-based access management. - **Monitoring & Continuous Improvement**: Ensures GenAI systems can be tested before moving to production and periodically monitored post-production. ## Asset Registries: Read more on different registries below: - [Data Inventory](table-registry/) - [LLMs and Models](model-catalog/) - [Prompts](prompts/) - [RAGs](rags/) - [End-to-End Pipeline Assets](pipelines/) ## How Platform Helps in Managing Inventories? The platform offers extensive capabilities to streamline onboarding and efficiently manage assets. - Multiple registries are available to centralize and manage smaller, reusable components of the pipeline. - Customized groups for creating assets within a registry. - Permissible Purpose Tracking that enables automatic validation to ensure components are used only for authorized purposes. - Flexible and granular access management. - Change History tracking for registered assets. - Lineage Tracking of registered assets. - Sharing of assets within and across teams. - Creating Custom Fields for the registry apart from the default ones. ## Metadata Tagging When any item is added to an entity - during the registration, various fields can be tagged to that object. Some basic fields are mandatory - for example: - Alias - A Python variable name to refer to the object by - Type - The data type that the object returns - Description - A free format field that can be used to describe the object being created. - Group - Useful to organize items, making them easier to search and find later - Permissible Purpose - A governance tracking mechanism to ensure items are used correctly - Location (of Data) - A data lake location where the data resides and can be fetched from - Training & Validation Data - Used when creating models New fields can be added to any of the registries to facilitate better inventory management in Settings > Fields & Screens section. Fields of various types can be added: - Short Text - Long Text - Date Time - File - Single Select - Multi Select - Multiple Files They can be marked as mandatory and customized with descriptions, placeholder values, default values, etc. and even be made mandatory to fill in. Values for fields can also be programmatically computed - with Python formulae. ### Data Type Data Types on the platform are useful to declare clear types that can be used for documentation. Data Types are very flexible on the platform. The types supported are: - Scalar Types: - Numerical - String - DateTime - Boolean - Array Types: - Array[Numerical] - Array[String] - Array[Array[DateTime]] - and other types can be created by mixing existing types ... - Struct Types: - Struct[decision: String, ranking: Numerical] - Struct[amt: Array[Numerical], flag: Boolean] - Struct[info: Map[Numerical,String],details: Struct[id: String,date: DateTime,age: Numerical]] - and other types can be created by mixing existing types ... - Map Types: - Map[Numerical, String] - Map[String, Numerical] - Map[Numerical, Array[Boolean]] - and other types can be created by mixing existing types ... --- # Global Functions Source: https://docs.genguardx.ai/register-and-refine/inventory-management/global-functions/ Markdown: https://docs.genguardx.ai/register-and-refine/inventory-management/global-functions/index.md Description: How Global Functions work in GGX — reusable analytical logic with inputs and outputs of any type, registered once and called across pipelines, RAGs, models, reports, and simulations. ## What is a Global Function? A **Global Function** lets you write a set of analytical operations **once** and run them many times with different objects as inputs — without rewriting the logic each time. It supports inputs and outputs of **any type** (DataFrames, dictionaries, plain Python values, and more) and does **not** require a predefined input/output schema. Because it is registered centrally, the same function can be called across the platform — inside GenAI Studio, Reports, Simulation Data Sources, or even other Global Functions — which is what makes it the platform's primary unit of reuse.
![Global Function anatomy: inputs of any type flow into reusable logic that returns an output of any type, and the same function is reused across GenAI Studio, Reports, Simulation Data Sources, and other Global Functions.](./global-function-concept.svg)
Write the logic once; call it with different inputs from anywhere on the platform.
## A worked example: masking card numbers Consider `mask_card`, a small utility that redacts card numbers in free text — keeping only the last four digits. It is the kind of operation many objects need: the [card-replacement assistant](../pipelines/#a-complete-example-a-card-replacement-assistant) calls it before logging a conversation, a RAG calls it to sanitise retrieved passages, and a compliance Report calls it on its output. Register it once, and all three reuse the same audited logic. ## Anatomy of a Global Function | Part | What it holds | Required? | |------|---------------|-----------| | **Definition / code** | The Python that performs the operation and returns a result. | | | **Input Arguments** | Each input the logic operates on, with its **Alias**, **Type**, optional flag, and default value. | Optional | | **Output Type** | The data type the function returns. | Optional — see note below | | **Resources** | Other registered Global Functions the logic can call. | Optional | | **Properties** | Description, Group, Approval Workflow. | Mostly required | | **Attributes** | — the Python variable name other objects call this function by. | | :::note[Output type is optional] Unlike most registered objects, a Global Function does **not** require an Output Type. If you leave it unset, the function may return **any** type — a DataFrame, a dictionary, a tuple, and so on. ::: ## Where Global Functions are used Once registered, a Global Function is available anywhere logic runs on the platform: Call it from a **Pipeline**, **Model**, or **RAG** as a registered Resource — shared pre/post-processing, formatting, or scoring helpers. Reuse the same logic when writing **Reports**, so report calculations match what runs in production. Prepare or transform the data that feeds a **bulk simulation**. Compose functions — a Global Function can call other registered Global Functions. ## Adding a Global Function to the registry The **Global Function Registry** is the central place where every registered utility lives, organised into customisable groups for easier tracking, monitoring, and creation. Click **Create** on the registry page, then work through the form: 1. **Name, Properties, and Attributes.** Give the function a clear name and description. Set the **Group** and **Approval Workflow** under Properties, and the **Output Type** (optional) and under Attributes. 2. **Input Arguments.** Define each argument with its **Alias**, **Type**, optional flag, and default value. 3. **Resources.** Select any other registered Global Functions the logic should be able to call. 4. **Definition source.** Write the Python that performs the operation and returns the result. Use **Test Code** to run it against sample inputs before saving. 5. **Additional Information.** Add notes or attach supporting documentation. 6. **Save.** Click **Save** to register. The function is saved as a **Draft** until it goes through approval. ## A complete example: `mask_card` This fills in every field for the `mask_card` utility introduced above.
Page fields for mask_card | Field | Value | |-------|-------| | Description | *"Redacts card numbers in free text, keeping only the last four digits. Reusable across pipelines, RAGs, and reports for safe logging and display."* | | Alias | `mask_card` | | Output Type | `str` | | Input Arguments | `text` (`str`, required) | | Resources | — |
```python title="mask_card — definition source" import re # Replace any 13–16 digit sequence with •••• and its last four digits def _mask(match): digits = re.sub(r"\D", "", match.group()) return "•••• " + digits[-4:] return re.sub(r"\b(?:\d[ -]?){13,16}\b", _mask, text) ``` :::note `text` is the Input Argument defined in Step 2, and the return value matches the declared **Output Type** `str`. ::: Once registered, call it by its **Alias** from any pipeline, RAG, model, or report: ```python # inside a pipeline's scoring logic, before logging the turn safe_message = mask_card(user_message) ```
## Testing a Global Function While writing the definition, click **Test Code** at the bottom of the editor to run the function against sample inputs without saving — confirm it returns what you expect and that the return value matches the declared **Output Type** (if you set one). Because functions are composed into other objects, they also get exercised end-to-end whenever a pipeline, RAG, or report that uses them is run or simulated. ## Capabilities unlocked by registration Registering a Global Function — rather than copy-pasting a helper into every script — is what turns it into a governed, reusable asset: | Capability | What you get | |------------|--------------| | **Reusability** | Call the same logic across pipelines, models, RAGs, reports, and simulations, with visibility through [Lineage Tracking](../../lineage-tracking/). | | **Change tracking** | Automatic recording of modifications with efficient version upgrades. | | **Auditable path to production** | A transparent, fully auditable journey from Draft through Approval to downstream use. | | **Better collaboration** | A shared base for continuous building and testing across teams. | --- # Model Catalog Source: https://docs.genguardx.ai/register-and-refine/inventory-management/model-catalog/ Markdown: https://docs.genguardx.ai/register-and-refine/inventory-management/model-catalog/index.md Description: How models work in GGX — API-based, Python-based, and custom-uploaded models, how to register one, the supported providers, how to test a model, and a worked example. ## What is a model? A **model** is a software program that uses algorithms or rules to make informed decisions, predictions, or generations from a set of inputs without being given explicit instructions for every scenario — ML models, lookup tables, if-else rules, and LLMs all qualify. A registered model in GGX typically includes: - **Model file** — stored weights, parameters, lookup tables, tensors, or other data needed to initialise the model. - **Scoring Logic** — the code that takes inputs and produces a prediction or generation.
![Model anatomy: typed inputs flow into scoring logic, which calls a model source — API provider, Python logic, or uploaded weights — and returns the model output.](./model-concept.svg)
Inputs flow into scoring logic, which calls a model source and returns the model's output.
## The three model types Every model registered in GGX is one of three types. Choosing the right one is the first real decision you make on the registration form — it determines where the model runs and how it is configured. **Best for** hosted foundation models — OpenAI, Anthropic, Google Vertex AI, Azure OpenAI, AWS Bedrock, Hugging Face Inference. - GGX calls the provider over HTTPS using your configured credentials. - You pick the **Model Provider** and the specific **Model** from the dropdown. - You do not upload any weights. **Best for** lightweight Python logic, rule-based models, or anything that runs inside GGX without an external file. - Write the model's logic directly in the Scoring Logic editor. - Use any libraries packaged with the platform. - No upload, no external provider. **Best for** trained or fine-tuned models you want to host inside GGX — scikit-learn, BERT and other NLP models, custom fine-tunes. - Upload the weights or model file. - Write Scoring Logic that loads the file and produces a prediction. - Runs entirely inside GGX with no external API call. ## Adding a model to the catalog The **Model Catalog** is the central place where every registered model lives, organised into customisable groups. From here you can track, monitor, test, and create new models. Click **Create** on the Model Catalog page, then work through the form: 1. **Name and description.** Give the model a clear name and a description of what it does and when to use it. 2. **Properties.** Set the **Group**, **Permissible Purpose**, **Approval Workflow**, **Ownership Type** (Proprietary, Open Source, Internal), and **Model Type** (for example, *LLM*). 3. **Alias.** A code-safe variable name pipelines use to refer to this model — lowercase with underscores, no spaces. 4. **Input Type.** Pick **API-Based**, **Python-Based**, or **Custom**. If API-Based, also pick the **Model Provider** and the specific **Model**. 5. **Output Type.** The data type the model returns, e.g. `dict[str, str]`. 6. **Input Arguments.** For each argument (typically `text`, `temperature`, `system_instruction`, etc.), set its **Alias**, **Type**, whether it is optional, and a default value. 7. **Resources and weights.** Attach any registered Global Functions or Prompts the scoring logic needs. Upload the model file under **Pipeline Model File** if the type is Custom. 8. **Scoring Logic.** Write the Python that initialises the model and produces a result. Use **Test Code** to validate it against sample input. 9. **Save.** Add notes or attach documentation under **Additional Information**, then click **Create**. The model is saved as a **Draft** until it goes through approval. :::note[Credentials live in Integrations, not in code] For API-Based models, authenticate using environment variables exposed through **Platform Integrations** (e.g. `GOOGLE_API_TOKEN`, `OPENAI_API_KEY`). Do not paste keys into Scoring Logic. :::
## Supported providers API-Based models can connect to any of the following providers. Each integration page covers the credentials and configuration the provider expects. | Provider | Use it for | |----------|------------| | [OpenAI](../../../integrations/llm-providers/openai/) | GPT family and OpenAI-hosted models. | | [Anthropic](../../../integrations/llm-providers/anthropic/) | Claude family. | | [AWS Bedrock](../../../integrations/llm-providers/aws-bedrock/) | Bedrock-hosted foundation models from multiple vendors. | | [Google Vertex AI](../../../integrations/llm-providers/gcp-vertexai/) | Gemini and Vertex-hosted models. | | [Azure AI](../../../integrations/llm-providers/azureai/) | Azure-hosted OpenAI and other Azure foundation models. | | [Hugging Face](../../../integrations/llm-providers/huggingface/) | Inference endpoints for open-source models. | ## Testing a model Testing a model means confirming three things: the scoring logic runs without error, it can reach its source (provider API, uploaded file, or in-platform code), and the output matches the **Output Type** you declared. GGX gives you two levels for this. ### Quick check — Test Code in the editor Inside the Model Catalog page, the **Test Code** button at the bottom-right of the Scoring Logic editor runs the model against sample inputs without saving. Use it during development to: - Confirm the API credentials configured in Integrations actually reach the provider. - Verify the return value matches the declared **Output Type** (e.g. `dict[str, str]`). - Sanity-check temperature, max-token, and system-instruction handling. - Catch import errors or missing dependencies before saving. ### Bulk simulation across a dataset A single Test Code call tells you the model *works*; a **Bulk Simulation** tells you how it behaves across many real cases. It runs every row of a dataset through the model and produces one output per record — useful for: - Spotting edge cases (empty input, very long input, non-English) a single test would miss. - Measuring quality across a representative sample before promoting to production. - Attaching the run as evidence in the model's risk-assessment evidence tab. ## A worked example: Gemini 2.0 Flash Registering Google's **Gemini 2.0 Flash** as an API-Based model. The form fields: | Field | Value | |-------|-------| | Name | Gemini 2.0 Flash | | Alias | `gemini_2_0_flash` | | Input Type | API-Based | | Model Provider | Google Vertex AI | | Output Type | `dict[str, str]` | | Input Arguments | `text` (String, required) · `temperature` (Numerical, optional, default `0`) · `system_instruction` (String, optional, default `None`) | The scoring logic authenticates with an environment-variable token and calls Vertex AI: ```python title="gemini_2_0_flash — scoring logic" import os from google import genai from google.genai import types client = genai.Client(api_key=os.getenv("GOOGLE_API_TOKEN")) config = types.GenerateContentConfig( temperature=temperature, seed=2025, system_instruction=system_instruction, ) response = client.models.generate_content( model="gemini-2.0-flash", contents=text, config=config, ) return {"response": response.text} ``` :::tip[Where the API token comes from] You never hard-code the key. The `os.getenv("GOOGLE_API_TOKEN")` value is supplied by connecting the provider on the **Settings → Platform Integrations** page — pick the provider card (here **Google Vertex AI**), add its credentials, and GGX exposes them to scoring logic as environment variables. If your provider is **not** one of the listed cards, open the **Advanced** tab to register a custom provider and define the environment variable yourself, then reference it the same way with `os.getenv(...)`. ::: Once saved, it can be called from any downstream pipeline: ```python reply = gemini_2_0_flash( text=user_prompt, temperature=0.7, system_instruction="You are a helpful assistant.", ) output_text = reply["response"] ``` ## Models as evaluators (LLM-as-a-judge) A registered model is not only the thing being tested — it can also be the thing that *does* the testing. An **LLM-as-a-judge** is simply a model whose scoring logic pairs an LLM with a **Prompt** to score another system's output against criteria, returning a structured score (for example, answer relevancy on a `0–4` scale) and a reason. Because a judge is a registered model, it is **reusable** across any [Report](../../../evaluate-and-approve/reporting/) and pre/post-production evaluation, and it can itself be **validated**: run it over a ground-truth dataset with [Bulk Simulation](../../../evaluate-and-approve/simulation/) and compare its verdicts to the known answers before you trust it. Refine its prompt — including via [Prompt Optimization](../../prompt-optimization/) — to adapt a generic judge to your use case. :::tip[Heuristics first, judges for nuance] Pair judges with cheap, deterministic checks. Rule-based heuristics (token, length, and cost thresholds; staleness; keyword failures; library-based PII detection) triage the obvious cases instantly; reserve LLM judges for the quality questions rules cannot answer. ::: ## Capabilities unlocked by registration Registering a model — rather than calling it from a one-off script — is what turns it into a governed, reusable asset: | Capability | What you get | |------------|--------------| | **Change tracking** | Every modification to a draft is snapshotted in Change History; approved versions are locked. | | **Purpose enforcement** | Automatic detection of Permissible Purpose violations when the model is used downstream. | | **Testing & evaluation** | Quick Test Code during development and Bulk Simulation across datasets before promoting. | | **Reusability** | Reuse across pipelines, with visibility through [Lineage Tracking](../../lineage-tracking/). | | **API fingerprinting** | External API connectivity is fingerprinted so changes upstream are detectable. | | **Auditable path to production** | A transparent, fully auditable journey from Draft through Approval to use in pipelines. | | **Executable artifacts** | Extract ready-to-productionise artifacts straight from the Catalog. | --- # Pipelines Source: https://docs.genguardx.ai/register-and-refine/inventory-management/pipelines/ Markdown: https://docs.genguardx.ai/register-and-refine/inventory-management/pipelines/index.md Description: How pipelines work in GGX — compose Models, RAGs, Prompts and Guardrails with orchestration logic. Covers pipeline types, anatomy, registration steps, testing, and risk assessment. ## What is a pipeline? A **pipeline** is how you turn individual building blocks into a working GenAI application in GGX. It wires reusable components — Models, RAGs, Prompts, Guardrails, and even other pipelines — together with a piece of orchestration code, then takes an input, runs it through that logic, and returns a generated or predicted output. :::tip[A simple way to picture it] Think of a pipeline as a **recipe**: the registered components are the ingredients, and the **orchestration logic** is the set of instructions that decides how they combine. ::: ## A worked example: an IVR assistant Consider an **IVR (Interactive Voice Response)** assistant that automates customer support calls. Rather than building it as one large block of logic, you compose it from two smaller, reusable sub-pipelines.
![Flow of an IVR pipeline: a caller message passes through Intent Classification then Response Generation to produce a spoken reply.](./ivr-pipeline-flow.svg)
The IVR pipeline orchestrates two reusable sub-pipelines in sequence.
The IVR pipeline simply orchestrates the two in order: 1. The caller's message arrives. 2. **Intent Classification** labels the intent. 3. **Response Generation** uses that intent to produce the final reply. Because each sub-pipeline is registered on its own, you can reuse it elsewhere — the same Intent Classification sub-pipeline could power a chat widget or an email-triage tool. :::note[End-to-end flow] 1. Caller says: *"I lost my credit card."* 2. Intent Classification (LLM + Prompt) returns intent: `Report Lost Card`. 3. Response Generation (LLM + Prompt + RAG) replies: *"I'm sorry to hear that. I've blocked your card. Would you like a replacement?"* ::: ## The three pipeline types Every pipeline in GGX is one of three types. Choosing the right one is the first real decision you make when registering — it determines how the pipeline handles memory and what shape its output takes. **Best for** conversations, assistants, and multi-turn support. - Keeps `history` and `context` across turns - Stays aware of an ongoing conversation - Output is fixed: `{"output": ..., "context": ...}` *Example:* a support chatbot that recalls earlier messages. **Best for** one-shot generative tasks — summarization, drafting, extraction. - No memory — each input is independent - Processes one input and is done - Output shape is whatever you define *Example:* a pipeline that summarizes a document in a single pass. **Best for** sorting an input into a known category — sentiment, intent, routing. - Returns one label from a set you predefine - Can run single-turn, or take prior `context` - Output is a fixed label, e.g. `positive` *Example:* a pipeline that labels a review as `positive` or `negative`. :::tip[Which should I choose?] Pick **Chat-Based** for a back-and-forth conversation, **Classification** when the answer must be one of a fixed set of labels, and **Free-Flow** for any other one-shot task. ::: ## Anatomy of a pipeline Regardless of type, every pipeline has the same three-part shape: it receives **inputs**, runs **scoring logic** that draws on **registered resources**, and returns an **output**.
![Pipeline anatomy: an input flows into scoring logic, which draws on registered resources, and returns an output.](./pipeline-anatomy.svg)
Inputs flow into scoring logic, which calls on registered resources to produce an output.
:::note For chat pipelines, the returned `context` is fed back in on the next turn — that is how the pipeline stays aware of an ongoing conversation. ::: ## Adding a pipeline to the registry The **Pipeline Registry** is the central place where all registered pipelines live, organized into customizable groups. From here you can track, monitor, test, and create pipelines. There are two ways to add one: Build the logic directly in GGX with the Python Function Editor. A transparent, white-box approach: every component is registered inside GGX and stitched together, so each can be independently tested, validated, and debugged. Integrate a pipeline already running elsewhere through its API. A black-box approach where the agent's internals are abstracted away and interaction is limited to invoking it and consuming its outputs. ## Registering a pipeline from scratch Click **Create** on the Pipeline Registry page, then work through the registration form: 1. **Name and description.** Give the pipeline a clear name, then a plain-English description of what it does, when to use it, and when not to — it is what teammates read when deciding whether to reuse it. Under **Add Additional Details**, set the **Group**, **Permissible Purpose**, and metadata like **Usecase Type** and **Task Type**. 2. **Alias.** A code-safe variable name other pipelines use to refer to this one — lowercase with underscores, no spaces. For our running example: `card_assistant`. 3. **Input type.** Choose how the logic is supplied: - **Python Function** — write the logic in the Scoring Logic editor on this page. - **External Agent** — connect to an agent built elsewhere. 4. **Pipeline, interaction, and context type.** These three move together: - **Pipeline Type** — pick **Chat-Based**, **Free-Flow**, or **Classification**. - **Interaction Type** — set automatically from Pipeline Type; for chat it is `TypedDict[{'role': str, 'content': str}]`. - **Context Type** — required for chat pipelines; the data type of the `context` carried between turns, e.g. `dict[str, str]`. 5. **Config, resources, and model file.** Attach what the logic needs: - **Add Config** — input arguments or configuration values, each with a type and default. - **Add Resources** — registered Models, Prompts, RAGs, Global Functions, Guardrails, or pipelines. - **Add Pipeline Model File** — a custom model or supporting file, if required. 6. **Write the scoring logic.** In the editor, write the Python that ties everything together and returns the result. Use **Format Code** to tidy it and **Test Code** to run it against sample input before saving. 7. **Finish registration.** Optionally add starting examples for human-in-the-loop testing and notes under **Additional Information**, then click **Create**. The pipeline is saved as a **Draft** until promoted. ## Variables available in scoring logic When you write scoring logic for a **chat pipeline**, several variables are provided automatically — you do not need to declare them. | Variable | Type | What it holds | |----------|------|---------------| | `user_message` | `str` | The current message from the user. | | `history` | `list[TypedDict[{'role': str, 'content': str}]]` | All previous messages, in standard OpenAI format. | | `context` | your **Context Type** | Information carried over from earlier turns. | | `cache` | `dict` | A store for intermediate, reusable objects so expensive work is not repeated. One cache per pipeline, shared across all executions. | :::note[About the output] The output of a chat pipeline is fixed as a dictionary — `{"output": string, "context": custom type}`. The `output` is the reply shown to the user; the `context` is whatever you want available on the next turn, typed by the **Context Type** from Step 4. ::: ## Testing your pipeline Once a pipeline is created, validate its behaviour before relying on it. There are three ways to test, in increasing order of thoroughness.
![Three escalating test levels: Quick Test, Interactive Test, and Bulk Simulation, in increasing order of thoroughness.](./pipeline-testing-levels.svg)
From a quick sanity check on one input to a full run over an entire dataset.
### Quick Test — during development A fast sanity check while you are still writing the logic; it runs the code without saving. 1. While creating or editing the pipeline, scroll to the **Code** section. 2. Click **Test Code** in the bottom-right corner of the editor. 3. Enter sample inputs and confirm the logic returns what you expect. ### Interactive Test — feedbacks after initial version is created Once saved, test the pipeline the way an end user would experience it. 1. Navigate to your saved pipeline. 2. Click **Run → Chat Session** in the top-right corner. 3. Enter sample messages to walk through the full conversation flow. 4. Confirm the outputs match the expected behaviour. :::note[Chat Session is for chat-based pipelines] **Chat Session** is available only for **chat-based** pipelines. For **free-flow** and **classification** pipelines, use **Test Code** instead — call the function with sample inputs to check the output. ::: ### Bulk Simulation — validation at scale before for approvals A single input tells you the pipeline *works*; a **bulk simulation** tells you how it behaves across many real cases. It runs every row of a dataset through the pipeline, producing one output per record. Use it to: - Spot edge cases and inconsistent outputs a single test would miss. - Measure quality across a representative dataset before promoting. - Attach the run as evidence in the pipeline's **Risk Assessment** tab. :::tip[When to use which] Reach for **Quick Test** while writing the logic, **Interactive Test** to feel the end-user experience, and a **Bulk Simulation** before promoting to production — the most thorough check of the three. ::: ## A complete example: a card replacement assistant This example fills in every field for a realistic **chat-based pipeline** that helps customers replace a lost card.
Page fields for card_assistant | Field | Value | |-------|-------| | Description | *"Conversational assistant that helps a customer report a lost card and request a replacement. Use for card-servicing chats; not for fraud disputes."* | | Alias | `card_assistant` | | Input Type | Python Function | | Pipeline Type | Chat Based Pipeline | | Context Type | `dict[str, str]` | | Resources | `kb` (a RAG over the card-policy knowledge base), `reply_prompt` (a Prompt), `chat_model` (an LLM) |
```python title="card_assistant — scoring logic" # Retrieve relevant card-policy passages for the user's message docs = kb.search(user_message, top_k=3) # (1)! # Fill the prompt with the retrieved policy and the conversation so far filled_prompt = reply_prompt( user_message=user_message, history=history, policy_docs=docs, ) # Generate the reply reply = chat_model(filled_prompt) # Track where we are in the flow so the next turn knows the # customer has already confirmed the card was lost updated_context = dict(context) updated_context["stage"] = "replacement_offered" # (2)! return {"output": reply, "context": updated_context} # (3)! ``` 1. `kb`, `reply_prompt`, and `chat_model` are **Resources** added in Step 5. `user_message` and `history` are provided automatically. 2. On the next turn, `context["stage"]` is available again — so the pipeline knows not to ask *"did you lose your card?"* twice. 3. Chat pipelines **must** return this exact `{"output", "context"}` shape. A **Free-Flow** pipeline is simpler — no history, no context, and you define the output shape. For example, scoring the urgency of a support ticket: ```python title="urgency_scorer — scoring logic" # `ticket_text` is a Config argument defined via Add Config filled_prompt = urgency_prompt(ticket=ticket_text) score = chat_model(filled_prompt) return {"urgency": score} ``` A **Classification** pipeline returns one label from a set you predefine at registration. For example, routing an incoming support message to the right team: ```python title="support_router — scoring logic" # Allowed labels are predefined when the pipeline is registered filled_prompt = routing_prompt(message=user_message) label = chat_model(filled_prompt) # one of: billing, technical, general return label ``` ## Connecting an external agent If a pipeline already runs in another environment, connect it to GGX instead of rebuilding it. On the New Pipeline page, set **Input Type** to **External Agent** and provide the agent's API connection. For example, an agent built in **Vertex AI Agent Playbooks** can be connected through its API. Integrations for common providers — **Vertex AI Agent Playbooks, AgentForce, Microsoft Copilot Studio, Vapi AI**, and others — are configured in the Integration Module, and new providers can be added by extending the framework. :::note Once connected, an external agent is tracked, tested, and monitored exactly like a pipeline built from scratch. ::: Browse available integrations
## Pipeline risk assessment Each pipeline has a dedicated **Risk Assessment** tab for documenting its assumptions, risks, and mitigations. Risks are recorded across several dimensions, and you can attach simulations as evidence of the testing performed. These are example dimensions — the set is configurable and can be auto-detected from a pipeline's configuration: | Dimension | What it captures | |-----------|------------------| | **Accuracy** | How reliably the pipeline produces correct results. | | **Stability** | How consistently it behaves across inputs and over time. | | **Ethics** | Fairness, bias, and responsible-use considerations. | | **Vulnerability** | Exposure to misuse, adversarial input, or failure modes. | ## Capabilities unlocked by registration Registering a pipeline — rather than running a loose script — is what turns it into a governed, reusable asset: | Capability | What you get | |------------|--------------| | **Change tracking** | Automatic recording of modifications, with efficient version upgrades. | | **Purpose enforcement** | Automatic detection of Permissible Purpose violations. | | **Testing & comparison** | Evaluate against other pipelines using custom and standardized validation kits. | | **Reusability** | Reuse across downstream applications, with visibility through Lineage Tracking. | | **Auditable path to production** | A transparent, fully auditable journey with easier production monitoring. | | **Human Integrated Testing** | Feedback logging and human-in-the-loop testing for chat-based pipelines. | | **Executable artifacts** | Extract ready-to-productionize artifacts straight from the Registry. | | **Better collaboration** | A shared base for continuous development and testing. | --- # Prompt Registry Source: https://docs.genguardx.ai/register-and-refine/inventory-management/prompts/ Markdown: https://docs.genguardx.ai/register-and-refine/inventory-management/prompts/index.md Description: How prompts work in GGX — the template, input arguments, and creation logic that make up a registered prompt, how to register one, how to test and improve it, and a worked example. ## What is a prompt? A **prompt** is a natural-language instruction given to a generative model to direct its response or produce a desired outcome. It may include questions, commands, contextual details, few-shot examples, or partial inputs for the model to complete or extend. In GGX, a prompt is more than the text you send — it is a **registered, versioned asset** made up of three parts that are stored and tracked together: - **Prompt Template** — the instruction text, with placeholders for any dynamic values. - **Input Arguments** — the typed inputs that fill those placeholders at runtime. - **Creation Logic** — Python that prepares the arguments and fills the template before the prompt is sent to a model.
![Prompt anatomy: a template with placeholders plus typed input arguments are merged by the creation logic into a filled prompt sent to a model.](./prompt-concept.svg)
The template and input arguments are merged by the creation logic into the prompt that is sent to a model.
:::note[When creation logic is trivial] If a prompt has no processing or formatting to do, the creation logic can simply return the template unchanged. ::: ## Anatomy of a prompt | Part | What it holds | Required? | |------|---------------|-----------| | **Prompt Template** | The instruction text. Placeholders use Python-style braces, e.g. `{customer_utterance}`. | | | **Input Arguments** | The typed values that replace placeholders. Each has an Alias, Type, optional flag, and default. | Optional for system-only prompts | | **Creation Logic** | A Python function that formats arguments and returns the filled template. | | | **Properties** | Description, Group, Permissible Purpose, Approval Workflow, Task Type, Prompt Type. | Mostly required — see below | | **Attributes** | Alias (the Python variable name pipelines call this prompt by). | | ## Adding a prompt to the registry The **Prompt Registry** is the central place where every registered prompt lives, organised into customisable groups. From here you can track, monitor, test, and create new prompts. Click **Create** on the Prompt Registry page, then work through the form: 1. **Name and description.** Give the prompt a clear name and a plain-English description of what it does, when to use it, and when not to — it is what teammates read when deciding whether to reuse it. 2. **Properties.** Set the **Group**, **Permissible Purpose**, **Approval Workflow**, **Task Type**, and **Prompt Type** (for example, *System Instruction*). 3. **Alias.** A code-safe variable name pipelines use to refer to this prompt — lowercase with underscores, no spaces. 4. **Resources.** Add any registered Models, Global Functions, or other assets the creation logic should be able to call. 5. **Input Arguments.** For each argument, set its **Alias**, **Type**, whether it is optional, and a default value. 6. **Prompt Template and Creation Logic.** Write the template with `{placeholder}` markers, then write the Creation Logic that fills them. Use **Test Code** to run it against sample input before saving. 7. **Save.** Optionally attach documentation or notes under **Additional Information**, then click **Save**. The prompt is saved as a **Draft** until it goes through approval.
## Testing and improving a prompt A prompt is only as good as the outputs it produces against the inputs you actually expect. GGX has dedicated features for prompt iteration *inside the Prompt Registry page itself*, plus the same Bulk Simulation that every other registered asset uses. ### Analyze Prompt The **Analyze Prompt** button at the bottom of the Prompt Template editor scores your prompt against a set of quality dimensions and returns an **estimated prompt score** (e.g. 75%). Alongside the score, GGX lists **Findings** — specific gaps in the prompt that, if addressed, would raise the score. For each dimension, GGX either lists one or more Findings with a short explanation, or shows *"Nothing was found"* if the dimension passes. The dimensions GGX currently evaluates include: - **Bias** — e.g. *"Minimal / No Fair Lending Instructions in the prompt"* — flagging missing guidance for the model to remain unbiased, fair, or avoid perpetuating stereotypes. - **Toxicity** — e.g. *"Minimal / No Code of Conduct instructions in the prompt"* — flagging missing guardrails against toxic, biased, or harmful output. - **Logical Progression and Coherence** — whether the prompt's instructions follow a clear, step-by-step structure. - **Examples Quality** — e.g. *"Limited Examples in the prompt"* or *"Limited Diversity in the prompt"* — flagging when there are too few examples, or when the examples don't reflect the range of real inputs the model will see. - **Clarity** — whether the prompt's instructions are unambiguous. - **Grammar** — e.g. *"Subject-Verb Disagreement"* — flagging grammatical issues in the prompt that could confuse the model. Use Analyze Prompt as your iteration loop — change the template, click Analyze again, watch the score move. ### Improve with AI The **Improve with AI** button uses the Findings surfaced by Analyze Prompt to generate suggested edits to the template — adding missing safety/bias guidance, sharpening unclear sections, restructuring for coherence. Review each suggestion, adapt it to your context, then re-run Analyze Prompt to confirm the score has improved before saving. ### Test the creation logic If your Creation Logic does anything non-trivial (formats a list of intents, fetches dynamic context, applies conditional logic), use **Test Code** at the bottom of the Creation Logic editor to run it against sample arguments and inspect the filled template before saving. ### Bulk simulation through a pipeline When the prompt is wired into a pipeline, **Bulk Simulation** runs that pipeline over an entire dataset of inputs and stores every filled prompt and model response — the right tool for measuring tone consistency, safety, output-format adherence, and edge-case behaviour at scale. ## A worked example: Customer Intent Classification Registering an intent-classification prompt for a banking assistant. The form fields: | Field | Value | |-------|-------| | Name | Customer Intent Classification | | Alias | `customer_intent_classification_prompt` | | Prompt Type | System Instruction | | Task Type | Classification | | Input Arguments | `user_message` (String, required) | The template uses two placeholders — the customer's query, and a list of intents that the creation logic formats from a Python list: ```text title="prompt template (excerpt)" You are a digital assistant for BankX. Classify the customer's query into one of the predefined intents. {list_of_intents} OUTPUT FORMAT: {"classified_intent": "str"} Customer query: {customer_utterance} ``` Creation Logic formats the intent definitions and fills both placeholders: ```python title="creation logic" intent_definitions = [ {"Intent": "ACTIVATE CARD", "Definition": "Request to activate a newly issued card", "Examples": ["How do I activate my new debit card?", "Activate my credit card now."]}, {"Intent": "BLOCK CARD", "Definition": "Request to block a lost, stolen, or compromised card", "Examples": ["Block my credit card immediately.", "I lost my debit card, can you block it?"]}, # ... ] def get_intent_info(data_list): """Format intent definitions into readable text.""" lines = [] for i, item in enumerate(data_list, 1): lines.append(f"#### {i}. {item['Intent']}") lines.append(f"- Definition: {item['Definition']}") for ex in item["Examples"]: lines.append(f" • {ex}") lines.append("") return "\n".join(lines) return prompt.format( customer_utterance=user_message, list_of_intents=get_intent_info(intent_definitions), ) ``` Once saved, a pipeline calls the prompt by its alias: ```python result = customer_intent_classification_prompt(user_message=user_input) intent = result["classified_intent"] ``` ## Capabilities unlocked by registration Registering a prompt — rather than hard-coding it in a script — is what turns it into a governed, reusable asset: | Capability | What you get | |------------|--------------| | **Change tracking** | Every modification to a draft is snapshotted in Change History; approved versions are locked. | | **Purpose enforcement** | Automatic detection of Permissible Purpose violations when the prompt is used downstream. | | **Testing & evaluation** | Analyze Prompt, Improve with AI, and Bulk Simulation through a pipeline. | | **Reusability** | Reuse across pipelines, with visibility through [Lineage Tracking](../../lineage-tracking/). | | **Auditable path to production** | A transparent, fully auditable journey from Draft through Approval to use in pipelines. | | **Executable artifacts** | Extract ready-to-productionise artifacts straight from the Registry. | --- # RAGs Source: https://docs.genguardx.ai/register-and-refine/inventory-management/rags/ Markdown: https://docs.genguardx.ai/register-and-refine/inventory-management/rags/index.md Description: How RAGs work in GGX — the knowledge source and retrieval logic that make up a registered RAG, how to register one, and what registering it unlocks. ## What is a RAG? **Retrieval-Augmented Generation (RAG)** is an AI approach that helps language models by integrating a retrieval mechanism that fetches relevant external information in real time. This information allows the model to generate more accurate, up-to-date, and context-aware responses beyond its pre-trained knowledge. A registered RAG in GGX has two main parts: - **Knowledge Source** — a repository of external information: documents, vector databases, knowledge graphs like Neo4j, or other structured/unstructured data sources. - **Retrieval Logic** — code that fetches the most relevant information from the knowledge source based on the provided inputs.
![RAG anatomy: a query flows into retrieval logic, which draws on a knowledge source — uploaded file, vector database, knowledge graph, or document store — and returns the retrieved information.](./rag-concept.svg)
A query flows into retrieval logic, which draws on a knowledge source and returns the retrieved information.
## A worked example: a card-policy knowledge base To make this concrete, picture `kb` — a RAG over a bank's **card-servicing policy** documents. It is the same `kb` that the [card-replacement assistant](../pipelines/#a-complete-example-a-card-replacement-assistant) calls whenever a customer asks to replace a lost card. On its own, `kb` does exactly one job: given a question, return the most relevant policy passages. 1. A query arrives — e.g. *"How long does a replacement card take?"* 2. The **retrieval logic** embeds the query and searches the **knowledge source** (a vector index built from the policy PDFs). 3. It returns the top-K passages — the grounding a downstream pipeline feeds to its model. Because `kb` is registered on its own, any pipeline can reuse it — the same knowledge base could ground a card-servicing chatbot, an email-triage tool, or an internal policy-search widget. ## Anatomy of a RAG | Part | What it holds | Required? | |------|---------------|-----------| | **Retrieval Logic** | Code that fetches relevant information from the knowledge source based on the inputs. | | | **Knowledge Source** | A repository of external information — documents, vector databases, knowledge graphs. | Uploaded for Custom; configured via API for API-Based | | **Input Arguments** | Typed inputs the retrieval logic operates on. Each has an Alias, Type, optional flag, and default value. | Optional | | **Properties** | Description, Group, Permissible Purpose, Approval Workflow. | Mostly required | | **Attributes** | Output Type and Alias (the Python variable name pipelines call this RAG by). | | ## The three retrieval types Every RAG registered in GGX is one of three types. The choice determines where the knowledge source lives and how the retrieval logic reaches it. Communicates with external knowledge sources like **Neo4j** or **vector databases** using APIs to retrieve information from outside environments. Lightweight Python logic using various libraries or rule-based retrieval systems. Leverages uploaded knowledge sources like **CSV** files or **vector indices** that GGX hosts as part of the RAG definition. ## Adding a RAG to the registry The **RAG Registry** is the central place where every registered RAG lives, organised into customisable groups. From here you can track, monitor, test, and create new RAGs. Click **Create** on the RAG Registry page, then work through the form: 1. **Name, Properties, and Attributes.** Give the RAG a clear name and description. Set the **Group**, **Permissible Purpose**, and **Approval Workflow** under Properties, and the **Output Type** and under Attributes. 2. **Input Arguments.** Define each argument with its **Alias**, **Type**, optional flag, and default value. 3. **Resources.** Select any registered Models, Global Functions, or Prompts the retrieval logic should be able to call. 4. **Input Type.** Pick **API-Based**, **Python-Based**, or **Custom**. 5. **Knowledge file and Retrieval Logic.** Upload the custom knowledge file if required, then write the retrieval code in the **Retrieval Logic** section. 6. **Additional Information.** Add notes or attach supporting documentation. 7. **Save.** Click **Save** to register. The RAG is saved as a **Draft** until it goes through approval. ## A complete example: the card-policy knowledge base This fills in every field for `kb`, the **Custom** RAG introduced above. It uploads a vector index built from the card-servicing policy documents and returns the passages most relevant to a query.
Page fields for kb | Field | Value | |-------|-------| | Description | *"Retrieves the most relevant passages from the bank's card-servicing policy documents. Use to ground card-servicing answers; not a source for fraud-dispute rules."* | | Alias | `kb` | | Input Type | Custom | | Output Type | `list[str]` | | Input Arguments | `query` (`str`, required), `top_k` (`int`, default `3`) | | Resources | `embedder` (an embedding Model) | | Knowledge file | `card_policy_index` — a vector index built from the policy PDFs |
```python title="kb — retrieval logic" # `query` and `top_k` are Input Arguments defined in Step 2. # `embedder` is a Resource; `card_policy_index` is the uploaded knowledge file. query_vector = embedder.embed(query) # (1)! hits = card_policy_index.search(query_vector, top_k=top_k) # (2)! # Return the passages most relevant to the query return [hit.text for hit in hits] # (3)! ``` 1. `embedder` is a registered Model added under **Resources**; `query` is provided as an Input Argument. 2. `card_policy_index` is the uploaded knowledge file; `top_k` defaults to `3` but the caller can override it. 3. The return value matches the **Output Type** `list[str]` — exactly what a pipeline receives when it calls `kb.search(...)`. An **API-Based** RAG reaches an external store — here a Neo4j knowledge graph — instead of an uploaded file: ```python title="graph_kb — retrieval logic" # `query` is an Input Argument; `graph` is a Resource holding the connection. cypher = build_cypher(query) records = graph.run(cypher, limit=top_k) return [r["passage"] for r in records] ```
## Testing a RAG A RAG is testable **on its own** — you do not need to wire it into a pipeline first. Validate retrieval quality independently with a **Quick Test** and a **Bulk Simulation**, then exercise it **end-to-end** inside any pipeline that uses it.
![Three escalating ways to test a RAG: a Quick Test on one query, a Bulk Simulation across a dataset of queries, and an end-to-end run inside a pipeline that uses the RAG.](./rag-testing-levels.svg)
From a single-query sanity check, to a full run over a dataset, to an end-to-end test inside a real pipeline.
### Quick Test — independent, while writing the logic A fast check on a single query without saving; it runs the retrieval logic against sample input so you can confirm it returns sensible passages. 1. While creating or editing the RAG, scroll to the **Retrieval Logic** section. 2. Click **Test Code** in the bottom-right corner of the editor. 3. Enter a sample `query` (and any other Input Arguments, like `top_k`) and confirm the returned chunks are relevant. ### Bulk Simulation — independent, at scale A single query tells you the RAG *runs*; a **bulk simulation** tells you how its retrieval behaves across many real questions. It runs an entire dataset of queries through the RAG and records one set of results per query — no pipeline required. Use it to: - Spot queries that retrieve irrelevant or empty passages a single test would miss. - Measure retrieval quality across a representative set of questions before approval. - Attach the run as evidence in the RAG's approval and risk review. Bulk Simulation is the same at-scale evaluation used across all registered assets, so the run, its dataset, and its results are logged and comparable just like a pipeline simulation. ### In a pipeline — end-to-end Once the RAG behaves on its own, test it in context. Any [pipeline](../pipelines/) that lists the RAG as a **Resource** exercises it as part of a full request — so you can see how retrieval quality shapes the final generated output, not just the raw chunks. This is where you confirm `kb` actually grounds the assistant's reply, rather than only returning plausible passages. :::tip[When to use which] Reach for **Quick Test** while writing the retrieval logic, **Bulk Simulation** to validate retrieval quality across many queries before approval, and a **pipeline run** to confirm the RAG holds up end-to-end inside a real application — the most thorough check of the three. ::: ## Capabilities unlocked by registration Registering a RAG — rather than calling a retriever from a loose script — is what turns it into a governed, reusable asset: | Capability | What you get | |------------|--------------| | **Change tracking** | Automatic recording of modifications with efficient version upgrades. | | **Purpose enforcement** | Automatic detection of Permissible Purpose violations. | | **Testing & evaluation** | Evaluate against other RAGs using custom and standardised validation kits. | | **Reusability** | Reuse across pipelines, with visibility through [Lineage Tracking](../../lineage-tracking/). | | **API fingerprinting** | External retrieval connectivity is fingerprinted so changes upstream are detectable. | | **Auditable path to production** | A transparent, fully auditable journey from Draft through Approval to use in pipelines. | | **Executable artifacts** | Extract ready-to-productionise artifacts straight from the Registry. | --- # Data Assets Source: https://docs.genguardx.ai/register-and-refine/inventory-management/table-registry/ Markdown: https://docs.genguardx.ai/register-and-refine/inventory-management/table-registry/index.md Description: Register data tables and quality checks in GGX so teams can track source data, fetch schemas, run data quality reports, audit changes, and reuse validation datasets. The Table Registry records the location, content, and structure of source data tables used for analytics. The table can be registered using one of the following methods: - **Connection to a Data Lake:** A direct link to a data lake allows specifying the location of the data file. The link must point to a valid **PySpark** data file. - **Table Upload:** Datasets with fewer rows can be uploaded directly in CSV or Excel format. > **Note:** Large or production-scale datasets should be registered through the **Data Lake connection** rather than uploaded through the browser. The data is then read directly from its source location (for example a cloud bucket or a PySpark/Parquet file), avoiding slow uploads while remaining viewable as a table. ## Managing Tables on the Platform The **Table Registry** organizes all the registered tables into customized groups at this centralized location and allows easier tracking, monitoring, and creating new ones. ### Registering a Table: - Click on **Create** button in Table Registry. - Fill in important details like **Name**, and **Attributes** (Alias, Group, Input Type, Location, Description). - Select an Input Type (Data Lake or Upload Data) and provide a data link or upload files accordingly. - Finally, click on the **Save** button to complete the registration. > **Note:** After registering the table, users can **Edit** and click on the **Fetch Columns** button to automatically load the table columns and their types. Once the table is registered, data quality can be evaluated through registered **Quality Checks** or it can be used for validation and testing. ## Benefits of Table Registration: - Automated **change history records** that tracks all the modifications to the tables. - **Track the lineage** of table usage in downstream applications. - Run **Quality Checks** on the tables. - Use tables for validation and testing in a **fully auditable** manner. - **Export tables** outside the platform when required with a single click. ## What is a Quality Check? The Quality Check enables the analysis of data and the creation of standard or custom reports based on registered tables in the Table Registry. It supports the generation of profiling metrics, descriptive statistics, invalid entry detection, outlier analysis, and other custom reports and metrics to assess data quality effectively before using the data for downstream tasks like running jobs. > **Note:** The Quality Check object currently supports linking only one table at a time, enabling the generation of multiple metrics and reports for a single table per analysis. ## Managing Quality Checks on the Platform: The **Quality Check Registry** organizes all the registered quality checks into customized groups at this centralized location and allows easier tracking, monitoring, and creating new ones. ### Registering a Quality Check: 1. Click on **Create** button in Quality Check registry. 2. Fill in important details like **Name**, **Attributes** (Data-Table, Group, Descriptions, Select Data-Columns). 3. **Add notes**, **attach documentation** if available in the **Additional Information** section. 4. Lastly, click on the **Save** button to complete the registration process. ## Benefits of Quality Check Registration: - **Analyze and monitor data** using standard and custom reports. - **Share data analysis and evaluations** with other team members. --- # Lineage Tracking Source: https://docs.genguardx.ai/register-and-refine/lineage-tracking/ Markdown: https://docs.genguardx.ai/register-and-refine/lineage-tracking/index.md Description: Use GGX lineage tracking to visualize object dependencies, identify upstream and downstream impact, support traceability, and understand how registered assets are used. Lineage tracking records and visualizes the complete lifecycle of an object and depicts the flow of execution, showing how objects interact, are transformed and executed. It captures its origins, building blocks used, transformations, and usages across the organization. ## Why Lineage Tracking is Important? - **Ensures Traceability** by providing a clear link between all artefacts, ensuring transparency in dependencies and usages. - **Easier Impact Assessment** for modifications. - Provides a clear view of all dependent objects that would be affected by changes. - Identifies necessary updates across dependencies to introduce changes. - **Enhances Collaboration** by providing visibility into how different teams and models interact within the pipeline. - **Ensures Compliance** with regulations like the **[EU AI Act (Article 13 & 14)](https://eur-lex.europa.eu/resource.html?uri=cellar:e0649735-a372-11eb-9585-01aa75ed71a1.0001.02/DOC_1&format=PDF)**, which mandates transparency and human oversight. - **Identifying redundant processes**. ## Lineage Structure on GGX Platform: A Lineage is created and automatically maintained by the platform as soon as an object is registered. The lineage graph which is present on the details page is a **horizontal-tree representation** that maps how an object is created, used, and executed within the pipeline. It consists of three key elements: - **Precedents (Inputs)**: Represent the sources or dependencies that contribute to the creation of an object. - **Dependents (Usage)**: Indicate other objects that rely on a given object. ![Lineage Example](./lineage-example.png) --- # Prompt Optimization Source: https://docs.genguardx.ai/register-and-refine/prompt-optimization/ Markdown: https://docs.genguardx.ai/register-and-refine/prompt-optimization/index.md Description: How GGX automates prompt optimization with Hill Climbing — iteratively refining a prompt, keeping only changes that improve a fixed evaluation, with full logs and one-click sync back to the pipeline. Prompt optimization is the process of **refining prompts** to get more accurate responses from an LLM. The goal is to make intent as clear as possible — but conveying every detail in one attempt is hard, so it takes **repeated testing and small adjustments** until the model delivers what you want. Doing that by hand — making a change, measuring its impact, and keeping a record of every experiment — is tedious and error-prone. GGX automates it with **Hill Climbing** experimentation. ## What is Hill Climbing? Hill Climbing is an optimization method that improves a solution **gradually**. It starts from an initial guess, makes small changes, and **keeps any change that produces a better result** — like climbing a hill by always stepping upward. It stops when no further improvement can be found.
![A rising performance curve: from an initial prompt at a low score, each small edit that raises the score is kept and steps up the curve, while edits that lower it are rejected — until the climb reaches the peak, the best prompt.](./hill-climb.svg)
Each kept edit steps the prompt further up the performance curve; edits that score worse are discarded.
## How it works A run holds the evaluation **fixed** and changes **only the prompt**, so any score difference is attributable to the prompt alone.
![The hill-climbing loop: Initialize, then Evaluate, Modify, Compare, and Update — keeping the new prompt as the baseline only if it scores better — then iterate, with evaluation holding the LLM, metrics, and dataset constant.](./optimization-loop.svg)
Initialize → Evaluate → Compare → Promote → Modify, looping until the target is reached.
1. **Initialize** — start with an initial prompt as the baseline. 2. **Evaluate** — score it against fixed components: a predefined **LLM with set hyperparameters**, **standardized metrics**, and a **consistent dataset**. 3. **Modify** — make a small, targeted change to the prompt. 4. **Compare** — re-run the evaluation and measure the change in performance. 5. **Promote** — if the new prompt scores better, it becomes the new baseline; if not, it is discarded. 6. **Iterate** — repeat until the desired performance is reached. ## What the capability gives you | Benefit | What you get | |---------|--------------| | **Full experiment log** | Every prompt update and its impact on performance is recorded automatically. | | **One-click sync-back** | Once the best-performing prompt is found, sync it back to the base pipeline. | | **Team collaboration** | Multiple members can work on the same optimization simultaneously. | | **Custom reports** | Configurable dashboards and reports track hill-climbing progress over time. | ## Running a Hill-Climbing task 1. Open the object's **Details** page. 2. Click **Run → Hill Climbing**. 3. **Provide a description** of the run. 4. Under **Dashboard Selection**, choose the dashboard to evaluate against. 5. Make the **changes to the prompts** that are part of the object. 6. Prepare the evaluation data under **Data Sources**. 7. Click **Run** and wait for the job to complete. 8. When it finishes, review the optimization progress in the **Dashboards**. :::tip[Start from your best prompt] Hill climbing builds on the baseline, so a stronger starting point reaches a higher peak in fewer iterations. ::: --- # GGX Sync Source: https://docs.genguardx.ai/register-and-refine/sync/ Markdown: https://docs.genguardx.ai/register-and-refine/sync/index.md Description: Use GGX Sync to declare, version, and synchronize prompts, models, RAGs, pipelines, global functions, and reports from Python code into GenGuardX. ## Overview The GGX Sync system enables you to programmatically manage and version-control your AI assets (Prompts, Models, RAGs, Pipelines, Global Functions, and Reports) directly from your development environment. Using Python decorators and a simple sync command, you can declare and synchronize your components to the GenGuardX platform. --- ## What Can Be Synced? GGX Sync supports six core component types: | Component | Purpose | Example Use Case | | --- | --- | --- | | **Prompt** | System instructions, templates, and persona definitions | Customer service chatbot instructions | | **Model** | LLM configurations and API integrations | Gemini, GPT, Claude model wrappers | | **RAG** | Retrieval-Augmented Generation systems | Database query systems, knowledge base retrieval | | **Pipeline** | End-to-end workflows combining multiple components | Complete chatbot with intent classification and response generation | | **Global Function** | Reusable utility functions | Data preprocessing, validation, formatting functions | | **Report** | Evaluation and monitoring reports with visualizations | Model performance dashboards, bias analysis reports | --- ## Getting Started ### Step 1: Installation ```bash pip install genguardx ``` ### Step 2: Initialize Connection Before syncing any components, you need to authenticate with your GenGuardX instance. ### Obtaining Your API Key 1. Log into your GenGuardX platform 2. Navigate to **Profile Section** → **Account Security** 3. Locate your **API Key** (format: `eyJI-XXXX-XXXX-XXXX-XXXX-XXXX-3fe3`) 4. Click **"How to use this key"** to view the initialization code ### Initialize in Your Code ```python import genguardx as ggx # Initialize connection to your GenGuardX instance ggx.init( api_url="https://your-ggx-instance.example.com", # Change for your instance api_key="your-api-key-here", ) ``` > Important: Replace the api_url if you're using a different GenGuardX instance (e.g., production, staging). The URL should point to your specific deployment. > --- ## Component Declaration & Sync Each component type uses a specific decorator pattern. Components can be declared and synced **independently** or used together in pipelines. ### 1. Prompts Prompts define system instructions, templates, or conversation guidelines. ### Decorator Syntax ```python import genguardx as ggx @ggx.Prompt.declare( name='My Prompt Name', # Optional: defaults to function name group='My Group', # Optional: organizational grouping task_type='Question Answering', # Optional: Classification | Summarization | etc. prompt_type='System Instruction', # Optional: User Prompt | Others prompt_elements=['Persona + Goal', 'Tone', 'Constraints'], # Optional: list of components ) def my_prompt_function(*, cache: dict = {}, prompt: str = "Your prompt text here"): """Docstring describing the prompt purpose.""" # -- BEGIN DEFINITION -- return prompt # Sync to platform ggx.sync(my_prompt_function) ``` --- ### 2. Models Models wrap LLM API calls with consistent interfaces and cost tracking. ### Decorator Syntax ```python import genguardx as ggx @ggx.Model.declare( name='My Model Name', # Optional: defaults to function name group='My Group', # Optional: organizational grouping ownership_type='Proprietary', # Optional: Proprietary | Open Source model_type='LLM', # Optional: LLM | Text Embedding | Guardrail | Judge Model | Others provider='openai', # Optional: openai | google | anthropic | etc. (for API-based models) model='gpt-4', # Optional: specific model identifier (required if provider is set) ) def my_model_function(text: str, temperature: float = 0.7, *, cache: dict = {}): """Docstring describing the model configuration.""" # -- BEGIN DEFINITION -- # Your model implementation return {"response": "...", "cost": "..."} # Sync to platform ggx.sync(my_model_function) ``` --- ### 3. RAGs (Retrieval-Augmented Generation) RAGs connect to knowledge bases, databases, or document stores to retrieve relevant context. ### Decorator Syntax ```python import pathlib import genguardx as ggx @ggx.Rag.declare( name='My RAG System', # Optional: defaults to function name group='My Group', # Optional: organizational grouping knowledge_base_format='Relational Database', # Optional: Vector Database | Graph Database | etc. provider=None, # Optional: provider identifier for API-based RAG systems ) def my_rag_function( query: str, *, cache: dict = {}, knowledge: pathlib.Path = pathlib.Path('data.db') ): """Docstring describing the RAG system.""" # -- BEGIN DEFINITION -- # Your RAG implementation return {"retrieved_data": "...", "query_used": "..."} # Sync to platform ggx.sync(my_rag_function) ``` --- ### 4. Pipelines Pipelines orchestrate multiple components into complete workflows. ### Decorator Syntax ```python import typing as t import genguardx as ggx @ggx.Pipeline.declare( name='My Pipeline', # Optional: defaults to function name group='My Group', # Optional: organizational grouping usecase_type='Question Answering', # Optional: Summarization | Translation task_type='Generative Responses', # Optional: Classification | Templated Responses | etc. impact='External Facing', # Optional: Internal Only | Internal - with external implications data_usage=['Customer Specific Data'], # Optional: list of data types pipeline_type='Chat based - OpenAI Spec', # Required: or 'Custom Return Type' ) def my_pipeline_function( user_message: str, history: list[t.TypedDict("T", {'role': str, 'content': str}, total=False)] = (), context: t.Optional[dict[str, str]] = None, *, cache: dict = {}, ): """Docstring describing the pipeline workflow.""" # -- BEGIN DEFINITION -- # Your pipeline implementation return {"output": "...", "context": "..."} # Sync to platform ggx.sync(my_pipeline_function) ``` --- ### 5. Global Functions Global Functions are reusable utility functions that can be referenced by other components. ### Decorator Syntax ```python import genguardx as ggx @ggx.GlobalFunction.declare( name='My Utility Function', # Optional: defaults to function name group='My Group', # Optional: organizational grouping ) def my_utility_function(input_text: str, max_length: int = 100, *, cache: dict = {}) -> str: """Docstring describing the utility function.""" # -- BEGIN DEFINITION -- # Your utility function implementation return input_text[:max_length] # Sync to platform ggx.sync(my_utility_function) ``` **Key Points**: - Global Functions can be called by other components (Pipelines, RAGs, Models, etc.) --- ### 6. Reports Reports generate evaluation dashboards with metrics and visualizations for monitoring AI systems. ### Decorator Syntax ```python import typing as t import genguardx as ggx @ggx.Report.declare( name='My Evaluation Report', # Optional: defaults to function name object_types=['PIPELINE', 'FOUNDATION_MODEL'], # Required: list of object types this report evaluates group='My Group', # Optional: organizational grouping risk_type='Accuracy', # Optional: Accuracy | Stability | Bias | Vulnerability | Toxicity | Others task_type='Classification', # Optional: Classification | Templated Responses | etc. risk_domain='Model Risk Management', # Optional: Model Risk Management | Fair Lending | Technology | Infosec | Others evaluation_methodology='Statistical / ML Algorithms', # Optional: LLM-as-a-Judge | Rule-based | Statistical / ML Algorithms | Others report_methodology='Custom methodology', # Optional: free text description parameters=[], # Optional: list of report parameters ) def my_evaluation_report(job: t.Any, data: t.Any, *, cache: dict = {}) -> t.Any: """Docstring describing the report purpose.""" # -- BEGIN DEFINITION -- # Your report logic here # Process data and return metrics return metrics_dict, processed_data # Sync to platform ggx.sync(my_evaluation_report) ``` **Key Points**: - Reports typically work with **ReportOutput**, **DataLogicExample**, and **AdditionalReportFigure** helper components - Helper components use the `report` parameter to associate with their parent report (e.g., `report='My Evaluation Report'`) - Report functions typically receive `job` and `data` parameters and return processed results - ReportOutput functions generate visualizations (Plotly figures, Pandas DataFrames, or HTML/Markdown strings) ## Key Concepts ### The `# -- BEGIN DEFINITION --` Anchor Every component function **must** include this special comment. Only code after this anchor is synced to the platform. ```python def my_function(): # Imports and setup code here (NOT synced) import os # -- BEGIN DEFINITION -- # Everything after this line IS synced result = process_data() return result ``` **Why?** This ensures only the core logic is synced, keeping your platform definitions clean and portable. ### The `cache` Parameter All component functions include a `cache: dict = {}` parameter. This is a **reserved keyword** that is: - Automatically removed from function inputs during sync (not exposed to end users) - Commonly used for storing database connections in RAG implementations **Example from RAG:** ```python if "sql_cursor_obj" not in cache: conn = sqlite3.connect(...) cache["sql_cursor_obj"] = conn.cursor() cursor = cache["sql_cursor_obj"] ``` ### Dependency Resolution When you sync a Pipeline, the system: 1. Inspects the function bytecode to find referenced objects 2. Checks if those objects have `_corridor_metadata` 3. Recursively syncs those dependencies by calling their respective sync handlers 4. Collects the version IDs returned from each sync operation 5. Includes those version IDs in the pipeline payload (e.g., `promptVersionIds`, `foundationModelVersionIds`) **Example:** ```python # Define components import genguardx as ggx @ggx.Prompt.declare(name='System Prompt') def system_prompt(*, cache: dict = {}, prompt: str = "You are a helpful assistant."): """System instruction for the assistant.""" # -- BEGIN DEFINITION -- return prompt @ggx.Model.declare(name='GPT-4', provider='openai', model='gpt-4', ownership_type='Proprietary', model_type='LLM') def gpt4(text: str, *, cache: dict = {}): """GPT-4 model wrapper.""" # -- BEGIN DEFINITION -- # Implementation return {"response": "..."} @ggx.Pipeline.declare(name='Q&A Pipeline', pipeline_type='Chat based - OpenAI Spec') def qa_pipeline( user_message: str, history: list[t.TypedDict("T", {'role': str, 'content': str}, total=False)] = (), context: t.Optional[dict[str, str]] = None, *, cache: dict = {} ): """Q&A pipeline combining prompt and model.""" # -- BEGIN DEFINITION -- prompt = system_prompt() response = gpt4(f"{prompt}\n{user_message}") return {"output": response["response"]} # Sync only the pipeline - dependencies sync automatically ggx.sync(qa_pipeline) ``` --- ## Independent vs. Pipeline Usage ### Independent Sync You can declare and sync components **without** using them in a pipeline: ```python # Declare a standalone prompt @ggx.Prompt.declare(name='Greeting Prompt', group='Standalone') def greeting_prompt(*, cache: dict = {}, prompt: str = "Hello! How can I help you today?"): """Greeting message prompt.""" # -- BEGIN DEFINITION -- return prompt # Sync it independently ggx.sync(greeting_prompt) ``` ### Pipeline Integration Or use components together in a pipeline (they'll sync automatically): ```python @ggx.Pipeline.declare(name='Greeter Bot', pipeline_type='Chat based - OpenAI Spec') def greeter_pipeline( user_message: str, history: list[t.TypedDict("T", {'role': str, 'content': str}, total=False)] = (), context: t.Optional[dict[str, str]] = None, *, cache: dict = {} ): """Simple greeter pipeline.""" # -- BEGIN DEFINITION -- prompt = greeting_prompt() # References the prompt return {"output": prompt} ggx.sync(greeter_pipeline) # Syncs both prompt and pipeline ``` --- ## Metadata Parameters Reference ### Common Parameters (All Components) - `name` (str, optional): Display name on platform. Defaults to function name. - `group` (str, optional): Organizational grouping. Must exist on platform. ### Prompt-Specific Parameters - `task_type` (str, optional): `'Classification'` | `'Question Answering'` | `'Information Extraction'` | `'Summarization'` | `'Code Generation'` | `'Transformation'` | `'Generation'` | `'Others'` - `prompt_type` (str, optional): `'System Instruction'` | `'User Prompt'` | `'Others'` - `prompt_elements` (list, optional): List of components: `['Persona + Goal', 'Tone', 'Task', 'Constraints', 'Context', 'Examples', 'Reasoning Steps', 'Output Format', 'Recap']` ### Model-Specific Parameters - `ownership_type` (str, optional): `'Open Source'` | `'Proprietary'` - `model_type` (str, optional): `'LLM'` | `'Text Embedding'` | `'Guardrail'` | `'Judge Model'` | `'Others'` - `provider` (str, optional): Provider identifier (e.g., `'openai'`, `'google'`, `'anthropic'`). If provided, model is treated as API-based. - `model` (str, optional): Specific model identifier (e.g., `'gpt-4'`, `'gemini-pro'`, `'claude-3-opus'`). Required if provider is specified. ### RAG-Specific Parameters - `knowledge_base_format` (str, optional): `'Vector Database'` | `'Graph Database'` | `'Relational Database'` | `'External Web-Search APIs'` | `'NoSQL'` | `'Document'` | `'Others'` - `provider` (str, optional): Provider identifier for API-based RAG systems. ### Pipeline-Specific Parameters - `usecase_type` (str, optional): `'Question Answering'` | `'Summarization'` | `'Translation'` - `task_type` (str, optional): `'Classification'` | `'Templated Responses'` | `'Generative Responses'` | `'Summarization'` | `'Others'` - `impact` (str, optional): `'External Facing'` | `'Internal Only'` | `'Internal - with external implications'` - `data_usage` (list, optional): List of data types: `['No Additional Data', 'General Public Data', 'Internal Policies/Data', 'Customer Specific Data']` - `pipeline_type` (str, required): `'Chat based - OpenAI Spec'` | `'Custom Return Type'` ### Global Function Parameters - `name` (str, optional): Display name on platform. Defaults to function name. - `group` (str, optional): Organizational grouping. Must exist on platform. ### Report-Specific Parameters - `object_types` (list, required): List of object types the report evaluates (e.g., `['PIPELINE', 'FOUNDATION_MODEL', 'PROMPT']`) - `name` (str, optional): Display name on platform. Defaults to function name. - `group` (str, optional): Organizational grouping. Must exist on platform. - `risk_type` (str, optional): `'Accuracy'` | `'Stability'` | `'Bias'` | `'Vulnerability'` | `'Toxicity'` | `'Others'` - `task_type` (str, optional): `'Classification'` | `'Templated Responses'` | `'Generative Responses'` | `'Summarization'` | `'Others'` - `risk_domain` (str, optional): `'Model Risk Management'` | `'Fair Lending'` | `'Technology'` | `'Infosec'` | `'Others'` - `evaluation_methodology` (str, optional): `'LLM-as-a-Judge'` | `'Rule-based'` | `'Statistical / ML Algorithms'` | `'Others'` - `report_methodology` (str, optional): Free-text description of the methodology - `parameters` (list, optional): List of parameter dictionaries for report execution --- ## Complete Workflow Example ```python import genguardx as ggx import typing as t # Step 1: Initialize ggx.init( api_url="https://your-ggx-instance.example.com", api_key="your-api-key-here", ) # Step 2: Declare components @ggx.Prompt.declare( name='FAQ Prompt', group='Support', task_type='Question Answering', prompt_type='System Instruction', ) def faq_prompt(*, cache: dict = {}, prompt: str = "Answer FAQs concisely and professionally."): """System prompt for FAQ bot.""" # -- BEGIN DEFINITION -- return prompt @ggx.Model.declare( name='GPT-3.5', group='Models', ownership_type='Proprietary', model_type='LLM', provider='openai', model='gpt-3.5-turbo', ) def gpt35(text: str, *, cache: dict = {}): """GPT-3.5 Turbo model wrapper.""" # -- BEGIN DEFINITION -- # Implementation here return {"response": "..."} @ggx.Pipeline.declare( name='FAQ Bot', group='Support', usecase_type='Question Answering', task_type='Templated Responses', impact='External Facing', data_usage=['General Public Data'], pipeline_type='Chat based - OpenAI Spec', ) def faq_pipeline( user_message: str, history: list[t.TypedDict("T", {'role': str, 'content': str}, total=False)] = (), context: t.Optional[dict[str, str]] = None, *, cache: dict = {} ): """Complete FAQ pipeline.""" # -- BEGIN DEFINITION -- prompt = faq_prompt() response = gpt35(f"{prompt}\n\nQuestion: {user_message}") return {"output": response["response"]} # Step 3: Sync (syncs all dependencies automatically) ggx.sync(faq_pipeline) ``` --- ## Best Practices ### ✅ Do's 1. **Always use the anchor comment**: `# -- BEGIN DEFINITION --` 2. **Validate groups exist**: Check your platform for valid group names before declaring 3. **Use type annotations**: Help the platform understand your data types 4. **Test locally first**: Run your functions before syncing to catch errors 5. **Use meaningful names**: Choose descriptive names for components 6. **Include the cache parameter**: Even if unused, include `cache: dict = {}` 7. **Write clear docstrings**: Document what each component does ### ❌ Don'ts 1. **Don't hardcode secrets**: Use environment variables for API keys and passwords 2. **Don't skip the cache parameter**: Required for all component functions 3. **Don't sync without testing**: Ensure functions work locally first 4. **Don't use undefined groups**: Verify group names exist on your platform instance 5. **Don't forget the anchor comment**: Missing anchor will cause sync errors --- ## Troubleshooting ### "Group not found" Warning **Problem**: Warning message: `├ [WARN] Group "YourGroupName" not found: Ignoring the group` **Cause**: The group specified in your declaration doesn't exist on the platform. **Solution**: - Check available groups in your platform UI - Create the group in platform settings first - Or remove the `group` parameter (it will default to None or the existing component's group) **Note**: This is a warning, not an error - the sync will still complete successfully. --- ### "Expected anchor comment" Error **Problem**: Error message: `AssertionError: Expected anchor comment # -- BEGIN DEFINITION -- to be present in source. Found None` **Cause**: Your function is missing the required `# -- BEGIN DEFINITION --` anchor comment. **Solution**: Add the anchor comment before your function's core logic: ```python import genguardx as ggx @ggx.Prompt.declare(name='My Prompt') def my_prompt(*, cache: dict = {}, prompt: str = "Hello"): """Docstring here""" # -- BEGIN DEFINITION -- # <-- Add this line return prompt ``` --- ### "Skipping X - not declared as a GGX object" Warning **Problem**: Warning message: `[WARN] Skipping "function_name" - as it is not declared as a GGX object` **Cause**: Your pipeline references a function that doesn't have a `@declare()` decorator. **Solution**: Add the appropriate decorator to the referenced function: ```python import genguardx as ggx # Before (causes warning) def helper_function(): return "result" # After (no warning) @ggx.Prompt.declare(name='Helper') def helper_function(*, cache: dict = {}, prompt: str = "result"): # -- BEGIN DEFINITION -- return prompt ``` **Note**: This is a warning - the sync completes, but the dependency won't be tracked or synced. --- ### Invalid Metadata Values Warning **Problem**: Warning messages like: - `├ [WARN] Task Type "YourTaskType" is invalid: Ignoring the Task Type` - `├ [WARN] Prompt Type "YourPromptType" is invalid: Ignoring the Prompt Type` **Cause**: The value provided for a metadata parameter doesn't match the allowed values. **Solution**: Check the Metadata Parameters Reference section above for the exact valid values. For example: - Task Type must be one of: `'Classification'`, `'Question Answering'`, `'Information Extraction'`, etc. - Prompt Type must be one of: `'System Instruction'`, `'User Prompt'`, `'Others'` **Note**: This is a warning - the sync completes, but the invalid metadata field is ignored. --- ### "Prompt must be a kwarg named 'prompt'" Error **Problem**: Error message: `ValueError: Prompt "your_prompt_name" must be a kwarg names: "prompt" which contains the prompt template as the default value` **Cause**: Prompt functions require a specific parameter format. **Solution**: Add `prompt: str = "your template"` as a keyword-only parameter: ```python import genguardx as ggx @ggx.Prompt.declare(name='My Prompt') def my_prompt(*, cache: dict = {}, prompt: str = "Your prompt text here"): """Docstring""" # -- BEGIN DEFINITION -- return prompt ``` --- ### Type Hint Error for dict/list **Problem**: Error message: `Unable to process type hint: '' as the dict has no inner type` **Cause**: Using bare `dict` or `list` types without specifying inner types. **Solution**: Use proper type hints with inner types: ```python # Bad context: dict = None # Good context: dict[str, str] = None context: t.Optional[dict[str, str]] = None # For lists history: list = () # Bad history: list[dict[str, str]] = () # Good ``` --- ## Additional Resources - **Platform Documentation**: Contact your platform administrator for instance-specific docs - **API Reference**: Available in your GenGuardX instance - **Support**: Contact your platform administrator or support team --- --- # Version Management Source: https://docs.genguardx.ai/register-and-refine/version-management/ Markdown: https://docs.genguardx.ai/register-and-refine/version-management/index.md Description: How GGX versions objects — the Draft → Approved → Clone lifecycle, snapshot-based Change History, reverting to any point, and how versions propagate downstream. Version management is the systematic tracking, organizing, and controlling of changes to an object. A new **version** is created every time an object's definition is first registered or later changed — so nothing is ever silently overwritten. ## Why it matters | Benefit | What it gives you | |---------|-------------------| | **Reproducibility** | Results and artifacts can be reproduced from the exact version that made them. | | **Collaboration** | Contributions from multiple users are managed without clobbering each other. | | **Auditability & compliance** | A clear, complete history of every change and version. | | **Rollback & recovery** | Revert to a stable version after a failure or an unintended edit. | | **Experiment tracking** | Compare different versions to see what changed and why. | | **Continuous improvement** | Track incremental changes and their impact over time. | ## The version lifecycle Every object follows the same path: it starts as an editable **Draft**, is sent for **Approval**, and becomes a **locked** version. Cloning a locked version starts the next draft — and the cycle repeats.
![The version lifecycle: a Draft (Version 1) is edited with every save snapshotted, sent for Approval, and locked as an immutable Approved version; cloning it creates Draft Version 2, which repeats the cycle.](./version-lifecycle.svg)
Draft → Approval → locked version → Clone → next draft — with edits snapshotted at every save.
1. **Draft (Version 1).** A Draft is created the moment an object is registered. Every modification is automatically logged to the **Change History** tab. 2. **Approval.** The Draft goes through the approval workflow. 3. **Approved & locked.** Once approved, the version is immutable — it can no longer be edited. 4. **Clone (Version 2).** To keep working, clone the approved version into a new Draft, which goes through the same cycle of changes and approvals. :::note[How versions propagate] Changes to a **Draft** propagate downstream **immediately**. An **approved** version, by contrast, is frozen — every downstream object that already uses it keeps that version until it is **explicitly upgraded** to the newer one. ::: ## Change History — a log of snapshots **Change History** is a structured log of every modification made to an object over time. It guarantees a clear audit trail and the ability to revert to any previous state.
![Change History as a log of snapshots: each save adds a row recording the action (Created or Changed) and what changed, with a version name. Any snapshot can be previewed, restored, restored as a copy, or named.](./change-history-log.svg)
Each save is captured as a named snapshot that can be previewed, restored, or restored as a copy.
Each entry in the log is a **snapshot** — an exact copy of the object at the moment it was saved: - The platform takes a snapshot **every time** the object is edited and saved. - A change can be undone by **restoring** the snapshot from that point in time. - Snapshots can be **named**, **previewed**, or restored **as a copy**. A comprehensive change history is what makes responsible AI governance possible — it: - **Enhances traceability** — track *when*, *why*, and *by whom* changes were made. - **Improves accountability** for responsible AI deployment and governance. - **Facilitates debugging** — identify and roll back problematic updates. - **Ensures compliance** with regulations such as the [EU AI Act (Article 12)](https://eur-lex.europa.eu/resource.html?uri=cellar:e0649735-a372-11eb-9585-01aa75ed71a1.0001.02/DOC_1&format=PDF), which mandates record-keeping for high-risk AI systems. - **Builds trust** through transparency into how an object evolves. ## Versioning approved items Once a Draft is approved it **cannot be modified**. To continue work, create a **clone** — a new version that copies the object's **Name, Alias, and Type**. This is deliberately different from editing a Draft: Draft changes propagate downstream instantly, whereas a new approved version is adopted downstream only when an object is explicitly updated to use it — so existing applications never shift underneath you. --- # GGX Monitoring Benchmarks Source: https://docs.genguardx.ai/technology/benchmarks/monitoring/ Markdown: https://docs.genguardx.ai/technology/benchmarks/monitoring/index.md Description: Benchmark GGX monitoring at scale across observability ingestion, heuristic triage, parallel LLM judges, automated routing, and targeted human review. GGX is designed to monitor large fleets of production AI applications. It can partition independent conversations, execute heuristic and LLM-based checks concurrently, and distribute the work across threads, processes, and multiple worker machines. This page presents GGX benchmarks for an example scenario. Use the results as estimates that you can adapt to your workload. :::caution Performance varies with machine specifications, judge complexity, LLM latency, rate limits, and other factors. We have documented the benchmark variables to make the results easier to interpret. ::: ## Scenario: Monitoring 10 Live AI agents Consider an organization deploying many AI applications for internal teams and external customers:
10 production AI agents
1,000 conversations
per agent/day
10,000 conversations
/day
50,000 turns/day
at ~5 turns per chat
![Ten AI agents each produce one thousand conversations per day, creating ten thousand conversations and around 50k turns for GGX to monitor.](./monitoring-volume.svg)
Assuming the agents are active for approximately **12 hours per day**, this averages about **14 conversations per minute** or **1.2 turns per second**. Real systems are bursty, so peak throughput may be higher. ### Assumptions - Conversation traces and metadata already exist in an observability system such as Datadog or Amazon CloudWatch. GGX reads from that source. - The judges use medium-sized LLMs, such as `gemini-3.5-flash`, `claude-4.5-haiku`, or `gpt-4o-mini`. - The judges have medium-complexity instructions of approximately 500 words. - Batch inference and prompt caching were not used. ### Infrastructure The tests ran on a machine with the following specifications: - Memory: 8 GB - Number of CPUs: 4 - LLM latency: This varies by provider. In the benchmark setup, a simple `hi` request returned in approximately 400 ms. ## Benchmark Measurements The benchmark ran five judge types for each record: - Toxicity - Answer Relevancy - Completeness - PII detection on the input - PII detection on the response The values below are the measured wall-clock results using Gemini 3.5 Flash. Increasing the thread count from 8 to 32 (4× more threads) produced an approximate **2–3x speedup** for batches of 100 records or more. For 100 records, this reduces the workload from up to 13 groups of 8 concurrent records to 4 groups of up to 32 concurrent records. | # Threads | Records | Time | Time/record | Throughput | | ---------- | ------: | ------------------------: | --------------------: | ---------------: | | 8 threads | 1 | 1.5 sec | 1.520 sec | 0.66 records/sec | | 8 threads | 10 | 6 sec | 0.632 sec | 1.58 records/sec | | 8 threads | 100 | 39 sec | 0.390 sec | 2.56 records/sec | | 8 threads | 1,000 | 381 sec | 0.381 sec | 2.62 records/sec | | 8 threads | 10,000 | 3,182 sec | 0.318 sec | 3.14 records/sec | | 32 threads | 1 | 1.5 sec | 1.549 sec | 0.65 records/sec | | 32 threads | 10 | 5 sec | 0.551 sec | 1.82 records/sec | | 32 threads | 100 | 16 sec | 0.160 sec | 6.23 records/sec | | 32 threads | 1,000 | 135 sec | 0.135 sec | 7.40 records/sec | | 32 threads | 10,000 | 1,236 sec | 0.124 sec | 8.09 records/sec | Across the measured batches of 100 to 10,000 records, the aggregate wall-clock improvement is approximately **2.6x**. ### Judge breakdown The following chart breaks down the 100-record measurement by judge. ## The Monitoring Workflow The preceding numbers cover 10,000 records evaluated by five judges each. Because not every record requires five judges in a real-world workflow, the following section extrapolates the results to a more representative monitoring scenario and its business impact. GGX turns raw production conversations into actionable assignments in three processing stages: **Heuristic tests**, **LLM-based tests**, and **Assignment**.
![GGX monitoring workflow: observability data is acquired in about one minute, heuristic tests assign clear problems, parallel LLM judges assess the remaining conversations, and GGX routes the results to developers, ground truth, or human review.](./monitoring-workflow.svg)
Deterministic checks handle clear cases first; parallel LLM judges add nuance only where rules are insufficient.
| Stage | Workload | What GGX does | | ------------------ | ----------------- | ------------------------------------------------------------------------------------------------------- | | **1. Acquire** | 10,000 chats/day | Fetches new traces and metadata from observability storage. | | **2a. Heuristics** | About 3,000 chats | Runs deterministic checks for clear policy, safety, cost, latency, format, or known failure conditions. | | **2b. LLM judges** | About 7,000 chats | Runs 4–5 independent judges only where heuristics are insufficient. | | **3. Assignment** | 10,000 outcomes | Routes each conversation to the appropriate operational destination. |
The benchmark uses three final assignments:
Route the finding to the responsible developer or team. GGX can create or link a Jira ticket according to the organization's routing and deduplication policy. Add the accepted conversation to ground truth managed through the GGX [Table Registry](../../../register-and-refine/inventory-management/table-registry/). Send the ambiguous conversation to a human reviewer in a GGX [Annotation Queue](../../../deploy-and-monitor/annotation-queues/).
## Business Impact Conversation-level monitoring is naturally parallel: one conversation can usually be evaluated independently of another. GGX partitions work by tenant, agent, time window, or conversation, then scales each execution stage independently. The slowest part of the monitoring workflow is typically the LLM-as-judge calls. Because these are network calls, they generally place a low burden on CPU and memory. Work can be divided among local execution threads/processes and scaled horizontally across worker machines. The useful concurrency is bounded by the allocated infrastructure and downstream LLM provider limits. ### Projected judgment time The judge stage is expected to dominate wall-clock time because GGX adds less than 5% overhead to LLM-judge execution time. Under the benchmark assumptions: - Heuristics assign **30%**, or 3,000 conversations, without an LLM. - The remaining **7,000 conversations** reach the judge stage. - Running **five judges** creates **35,000 judge calls**. For this estimate, we conservatively assume approximately **10 concurrent LLM calls**. Actual concurrency depends on the LLM provider's quota limits. | Judge profile | Time | Throughput | Projected time | | ----------------- | -----: | -----------: | -------------: | | Lower complexity | 200 ms | 50 calls/sec | **~ 12 min** | | Higher complexity | 1 sec | 10 calls/sec | **~ 50 min** | :::note[Why 200 ms and 1 second?] The benchmark models approximately **200 ms** for a lower-complexity judge and **1 second** for a higher-complexity judge. These are test profiles, not universal model guarantees. Real judge calls can be faster or slower depending on the provider, region, prompt and response lengths, model, throttling, and retry behavior. ::: ### Reducing the human-review burden GGX automates the high-confidence decisions and reserves human attention for conversations classified as **Unsure**. This reduces review volume without removing human oversight from ambiguous or high-risk cases. At 2 minutes per conversation, manually reviewing all 10,000 daily conversations would require **20,000 reviewer-minutes**, or about **333 reviewer-hours per day**. With GGX, human effort is proportional to the final uncertainty rate rather than total production volume: | Uncertain chats | Records/day | Reviewer time/day | | --------------: | -----------: | -----------------: | | 1% | 100 | 3 hr 20 min | | 5% | 500 | 16 hr 40 min | | 10% | 1,000 | 33 hr 20 min | The uncertainty rate is a result to measure, not a target to force downward. A lower escalation rate is only useful when heuristic and judge assignments continue to meet validated quality thresholds. The automated feedback loop converts monitoring data into ground truth and tracked issues, which can reduce uncertainty over time. Judges can use prior ground truth and identified issues as context for refinement. Over time, this can help align judge behavior with reviewer expectations. --- # Self-Hosted GGX Source: https://docs.genguardx.ai/technology/self-hosting/ Markdown: https://docs.genguardx.ai/technology/self-hosting/index.md Description: Install, configure, scale, back up, harden, and operate self-hosted GGX instances across Kubernetes, Terraform, cloud, Docker, and manual deployment options. :::note[Pipeline Hosting] For guides on how the analytics and pipelines written in GGX can be deployed to Production - refer to the [Direct to Production](../../deploy-and-monitor/direct-to-production/) guide. ::: Guides that cover the installation, configuration, and scaling of Self-Hosted GGX instances for analytical use. - [Minimum Requirements](installation/minimum-requirements/) - Installing on your own infrastructure - [Kubernetes](installation/kubernetes/) - [Terraform](installation/terraform/) - [Amazon Web Services (AWS)](installation/aws/) - [Microsoft Azure](installation/azure/) - [Google Cloud Platform (GCP)](installation/gcp/) - [Docker-based](installation/docker-based/) - [Manual](installation/manual/) - Configurations: How to configure your self-hosted instance of GGX - [SSO Integration - Microsoft AD, Okta, Google Workspace, etc.](configurations/saml/) - [RDBMS - Oracle, MS SQL Server, Postgres, etc.](configurations/database/) - [Web Servers - Nginx, Apache, etc.](configurations/web-servers/) - [Integrating packages - Wheelhouse, Artifactory, etc.](configurations/packages/) - [Automated Approval Steps - Jenkins, ServiceNow, JIRA, etc.](configurations/approvals/) - [Data Lakes - HDFS, Hive, Snowflake, etc.](configurations/datalake/) - [Notifications - Email, Slack, Teams, etc.](configurations/notifications/) - [Process Management - Systemd, Supervisor, etc.](configurations/process-management/) - Scaling to 100s and 1000s of users - [Concurrency - Increasing number of parallel runs](scaling/concurrency/) - [Scaling to number of users](scaling/scalability/) - [Backup Management](scaling/backups/) - [Hardening your GGX instance](hardening/) ## Architectural Overview The GGX analytical layer lets analysts test and validate their logic and get the required approvals and compliance checks. The production layer is NOT described here because GGX is isolated from the production side. GGX is divided into various components to keep it modular and enable easy scaling for cloud-based deployments and also to manage high loads without much change. Each of the components can be installed on separate machines or any subset can be installed in the same machine. The components are divided into: - **Web Application Server**: The web application server for the analytical UI of the platform - **API Server**: The API for business logic - **API - Celery worker**: The worker for asynchronous API tasks - **Spark - Celery worker**: The worker for asynchronous spark tasks - **Jupyter Notebook**: The Jupyter Notebook server for free-form analytical use - **File Management**: The file management server to manage files - **Metadata Database (SQL RDBMS)**: The database with all metadata provided in the Web Application - **Authentication Provider**: The identity and auth provider for access and permissions - **Proxy / Load Balancers**: Load Balancers / Proxies to simplify the install Here is a typical network diagram of how the installation would look like: ![Network Diagram](./ggx-network-diagram.excalidraw.svg) --- # CI CD Integrations Source: https://docs.genguardx.ai/technology/self-hosting/configurations/approvals/ Markdown: https://docs.genguardx.ai/technology/self-hosting/configurations/approvals/index.md Description: Integrate GGX approval workflows with external CI/CD and review systems by implementing custom approval handlers and exchanging review actions through APIs. There are scenarios when approval of an object is not confined to the reviewers registered on the platform. The client may have an external application, where other stakeholders are already onboarded and would want to review objects on the platform. GGX provides the ability to interact with such 3rd party applications, using API calls, and integrating their review. This can be accomplished by defining a custom handler that uses the base approval handler class exposed by GGX. ## Handler class The user would need to define a `CustomApprovalHandler` which would inherit GGX's base approval handler class: `corridor_api.config.handlers.ApprovalHandler` The logic for approval would be defined by the following method(s) inside `CustomApprovalHandler`. - `send_action(review, action)` - `receive_action(review_id, payload)` ## Example The example focuses on a model, which needs to be sent for review. We can configure what information we want to send to the tool (in `send_action`). The tool will expose an endpoint that would take that information, create a model entry on their side, carry out the necessary approval process, and send the feedback to us (in `receive_action`). `reviewId` is the key and would be used for communications. ```python from corridor_api.config.handlers import ApprovalHandler class CustomApprovalHandler(ApprovalHandler): name = 'external_tool' def send_action(self, review, action): ''' :param review: review_object :param action: Action taken by the user on the platform for the given review - 'Request Approval' - 'Resubmit' - 'Cancel' - 'Remind' - 'Edit' (if the object is in the 'Pending Approval' state and it is edited) ''' url = 'http://externaltool.example.com/cp_review/' # assuming the 3rd party app is running on PORT: 7006 review_id = review.id object_ = review.object # GGX object from corridor import Model # if we need to restrict the 3rd party approvals to Models only if not isinstance(object_, Model): raise NotImplementedError(f'{type(object_)} is not expected to be used with "{self.name}" tool!!!') json_info = { 'modelId': object_.parent_id, 'modelName': object_.name, 'modelVersion': object_.version, 'modelVersionId': object_.id, 'modelGroup': object_.group, 'createdBy': object_.created_by, 'reviewId': review_id, 'responsibilityId': review.responsibility.id, 'responsibilityName': review.responsibility.name, 'action': action, 'comment': review.comment, } headers = {} # any headers can be configured (optional) res = requests.post(url + str(review_id), json=json_info, headers=headers) return {'status': res.status_code} def receive_action(self, review_id, payload): ''' :param review_id: id corresponding to the review object (this is the same id that was sent by CP when requesting the review) :param payload: payload expects 2 kwargs - 'action': one of 'Accept'/'Need Info'/'Need Changes'/'Reject'/'Comment' - 'comment': any comment which the external_tool's reviewer makes :return: dictionary with `action` and `comment` for the review with id: `review_id` ''' # The external app needs to do a POST call with `action` and `comment` as part of the payload. # the endpoint would look like below (assuming corridor-api is running on port 5000): # `http://localhost:5000/api/v1/models/review/<>/external` action = payload.get('action') comment = payload.get('comment') # do some processing, if required comment = 'No comment' if comment is None else comment return {'action': action, 'comment': comment} ``` ## Configurations Approval handler-related configurations need to be set in `api_config.py` along with other configurations. (assuming the `CustomApprovalHandler` class is defined in the file `custom_approval_handler.py`). ```python THIRD_PARTY_APPROVALS = { 'external_tool': { 'handler': 'custom_approval_handler.CustomApprovalHandler', }, } ``` --- # Common Configs Source: https://docs.genguardx.ai/technology/self-hosting/configurations/common-configs/ Markdown: https://docs.genguardx.ai/technology/self-hosting/configurations/common-configs/index.md Description: Configure common self-hosted GGX settings for application services, API behavior, storage, authentication, workers, notebooks, logging, and deployment environments. After the installation, there are some configurations that need to be configured for each of the components. This section describes all the available configurations needed for each component to work correctly. Using these configurations the setup can be tweaked as needed. The configurations for each component needs to be present in the corresponding config file: - `api_config.py`: For the API Server and Celery Tasks - `app_config.py`: For the Web Application Server - `jupyterhub_config.py`: For the Jupyter Notebook (Jupyterhub) - `jupyter_notebook_config.py`: For the Jupyter Notebook (Notebook server) ## Configuration using Environment Variables It is possible to configure the platform using environment variables instead of (or in combination with) config files mentioned above. This might be more convenient in a cloud based deployment setting or while using a centralized secret management system (like Hashicorp Vault). When working with secret management systems, configurations could be loaded as environment variables during deployment of the platform. When configuring API/APP component via environment variables, prepend the configuration key with `CORRIDOR_`. Below are some examples for some standard data types, | Setting in `api_config.py` | Environment Variable Equivalent | | ------------------------------------------------- | ----------------------------------------------------------------- | | `LICENSE_KEY = xxxxxxx` | `export CORRIDOR_LICENSE_KEY=xxxxxxx` | | `WORKER_PROCESSES = 1` | `export CORRIDOR_WORKER_PROCESSES=1` | | `REQUIRE_SIMULATION = False` | `export CORRIDOR_REQUIRE_SIMULATION=false` | | `WORKER_QUEUES = ['api', 'spark', 'quick_spark']` | `export CORRIDOR_WORKER_QUEUES="['api', 'spark', 'quick_spark']"` | ## API The API Configurations help in controlling how the API Server and Celery workers behave. Some of the commonly used configurations are: - `LICENSE_KEY`: The GGX license key to use to enable the application - `API_KEYS`: The API keys to accept requests from - `SQLALCHEMY_DATABASE_URI`: The Database URI to connect to for the Metadata Database - `FS_URI`: The FileSystem URI to connect to for File Management ## Web Application These configurations help in controlling how the Web Application Server behaves. Some of the commonly used configurations are: - `SECRET_KEY`: Ensure a unique secret key for your setup is used - `REST_API_SERVER_URL`: The URL of the API Server for business login and metadata - `REST_API_KEY`: The API Key to use when connecting to the API Server - `NOTEBOOK_CONFIGS__link`: URL to a notebook solution ## Jupyter The Jupyter configurations are divided into 2 sections: jupyterhub and jupyter-notebook configurations. ## JupyterHub Configurations The configurations used by the GGX are the same as the standard [Jupyter Hub configurations](https://jupyter-notebook.readthedocs.io/en/stable/configuring/config_overview.html). Some of the commonly used configurations are: - `c.JupyterHub.bind_url`: The URL to host JupyterHub on - `c.Authenticator.auth_api_url`: The GGX Web Application Server (When using the GGX Authentication) - `c.Spawner.env_keep`: And environment variables to be kept when spawning the user jupyter-notebooks - `c.Authenticator.auth_api_url`: The API for the Authentication. The URL of the Web Application Server. There are also additional env variables needed by the GGX Python package: - `os.environ['CORRIDOR_API_URL']`: The GGX API Server URL - `os.environ['CORRIDOR_API_KEY']`: The GGX API Key to use (if set) --- # Metadata: Database Setup Source: https://docs.genguardx.ai/technology/self-hosting/configurations/database/ Markdown: https://docs.genguardx.ai/technology/self-hosting/configurations/database/index.md Description: Configure GGX metadata database connections for PostgreSQL, Oracle, SQL Server, and related self-hosted database settings, migrations, and operational requirements. This section describes the different database options available for the Metadata Database. Choosing the right database is critical, as it will impact the usability of the API and Web Application of the users. The metadata database stores all the information entered into the Web Application in a structured so it can be efficiently used by other components. The platform currently supports the following databases: - Oracle Database - Industry-standard - SQLite - Meant for testing only - MS SQL - Enterprise-level solution - PostgreSQL - Open-source and reliable ## POSTGRESQL To use postgresql, use the following configurations: SQLALCHEMY_DATABASE_URI = 'postgresql://:@/' ## SQLite Supported versions: sqlite 3+ Sqlite is an easy to setup database which is file-based. To use sqlite with GGX, set the following configuration: SQLALCHEMY_DATABASE_URI = 'sqlite:/// ## Oracle DB Supported versions: Oracle DB 19+ To use oracle, use the following configurations: SQLALCHEMY_DATABASE_URI = 'oracle://:@/' ## MS SQL Supported versions: SQL Server 2016+ This required the the unixODBC devel libraries (`yum install unixODBC-devel`) and [SQL Server ODBC driver](https://docs.microsoft.com/en-us/sql/connect/odbc/linux-mac/installing-the-microsoft-odbc-driver-for-sql-server) to be installed. To use mssql, use the following configurations: SQLALCHEMY_DATABASE_URI = 'mssql+pyodbc://:@/?driver=ODBC+Driver+17+for+SQL+Server' --- # Datalake Integration Source: https://docs.genguardx.ai/technology/self-hosting/configurations/datalake/ Markdown: https://docs.genguardx.ai/technology/self-hosting/configurations/datalake/index.md Description: Configure GGX data lake access for self-hosted deployments so jobs, reports, and analytics can read and write approved data locations. GGX provides the ability to connect to different kinds of data lakes which could have data saved as `parquet`, `orc`, `avro`, `hive tables` or any in any other format. The user could define a custom file handler that would have the logic to read/write to/from the data lake. ## Example The example focuses on creating a data source handler to read from hive tables. The user needs to inherit the base class: `DataSourceHandler` and define the functions: - `read_from_location` - `write_to_location` ```python from corridor_api.config.handlers import DataSourceHandler class HiveTable(DataSourceHandler): """ Consider a case where every data table is a table in the Hive metastore. The table `location` identifier is the table name. """ name = 'hive' write_format = 'parquet' def read_from_location(self, location, nrow=None): try: import findspark findspark.init() import pyspark except ImportError: import pyspark spark = pyspark.sql.SparkSession.builder.getOrCreate() data = spark.table(location) if nrow is not None: data = data.limit(nrow) return data def write_to_location(self, data, location, mode='error'): return data.write.format(self.write_format).saveAsTable(location) ``` ## Configuration Once the handler class is created, it can be set up in `api_config.py` as below: (assuming the handler is defined in a file called `hive_table_hander.py` alongside `api_config.py`) ```python LAKE_DATA_SOURCE_HANDLER = 'hive_table_handler.HiveTable' ``` --- # Email Notifications Source: https://docs.genguardx.ai/technology/self-hosting/configurations/notifications/ Markdown: https://docs.genguardx.ai/technology/self-hosting/configurations/notifications/index.md Description: Configure email notification settings for self-hosted GGX so approvals, alerts, reviews, and system events can reach the right users. GGX provides an option to send email notifications to the user on their registered email id when an event occurs on the platform (e.g. completion of simulation, approval request for an object). The different events for which notifications are triggered on the Platform: - Completion or failure of job run - Workflow Status change of an object - Review status changes or comments added during approval process - Sharing of an object :::note The user cannot customize which notifications will be sent as email. ::: ## Configuration The email notifications can be configured in `api_config.py` file with the parameter: `NOTIFICATION_PROVIDERS`. The value should be a dictionary with the key being `email`. The value for `email` should be a dictionary again with the email configuration details. The email configuration details include: - `from`: The email id from which the notifications are to be sent - `username`: Username corresponding to the email id - `password`: The password for the email id - `host`: The host of the SMTP server - `port`: The port number to use - `ssl`: Should ssl be used - `html`: Should the email be parsed as an HTML file ### Example ```python NOTIFICATION_PROVIDERS = { 'email': { 'from': 'user@example.com', 'username': 'user', 'password': 'password', 'host': 'smtp.server.com', 'port': 465, 'ssl': True, 'html': True, }, } ``` --- # Additional Libraries Source: https://docs.genguardx.ai/technology/self-hosting/configurations/packages/ Markdown: https://docs.genguardx.ai/technology/self-hosting/configurations/packages/index.md Description: Install and manage additional Python libraries for self-hosted GGX components, notebooks, workers, and analytical workloads. GGX provides an option to allow the users to be able to use additional libraries which are not available out of the box. ## Configuration There are 5 steps which have to be followed so that the user is able to validate the definition and run successful jobs on the platform. 1. Install the library in the virtual environment of API servers. 2. Install the library in the virtual environment of Worker-API servers. 3. Install the library in the virtual environment of Worker-Spark servers. 4. Install the library on the spark cluster. 5. Add the library to the Allowed Python Imports in the `Platform Settings` tab in the platform UI --- # Process Management Source: https://docs.genguardx.ai/technology/self-hosting/configurations/process-management/ Markdown: https://docs.genguardx.ai/technology/self-hosting/configurations/process-management/index.md Description: Run self-hosted GGX services under process managers such as systemd or Supervisor to manage startup, restarts, logs, and long-running components. For ease of maintenance and monitoring, it is recommended to use a process management tool to ensure the component daemons are running correctly. Using a process manager can simplify restarts, reboots, status-checks, logging, and configurations. The process management tools that are frequently used are Systemd, init-script, etc. In general, it is recommended to use the tool that the OS provides to handle these tasks. To simplify the installation, Supervisor can also be used, which provides application-level process management that is independent of the host setup. The tools available for process management require `sudo` access which might not be an option in some situations. GGX has its own process manager for each of the components, which uses daemonisation to handle long-running, background processes. ## Daemon Mode No additional configurations are required when running GGX processes using GGX-Daemons. We can use the below set of commands to start/stop/check_status. The logfile and pidfile for each process can be saved at custom locations by using the parameters: - logfile: location of the log file - pidfile: location of the pid file The default logfile directory is: `INSTALL_DIR/instances/INSTANCE/logs` The default pidfile directory is: `INSTALL_DIR/instances/INSTANCE/pids` To start the processes, we can do: ```sh INSTALL_DIR/venv-api/corridor-api daemon start INSTALL_DIR/venv-api/corridor-worker daemon start INSTALL_DIR/venv-app/corridor-app daemon start INSTALL_DIR/venv-jupyter/corridor-jupyter daemon start ``` To stop the processes, we can do: ```sh INSTALL_DIR/venv-api/corridor-api daemon stop INSTALL_DIR/venv-api/corridor-worker daemon stop INSTALL_DIR/venv-app/corridor-app daemon stop INSTALL_DIR/venv-jupyter/corridor-jupyter daemon stop ``` To check the status of the processes, we can do: ```sh INSTALL_DIR/venv-api/corridor-api daemon status INSTALL_DIR/venv-api/corridor-worker daemon status INSTALL_DIR/venv-app/corridor-app daemon status INSTALL_DIR/venv-jupyter/corridor-jupyter daemon status ``` ## Supervisor To use GGX with Supervisor, some useful configurations are: - `command`: The command to execute. Note, if 2 commands need to be executed, use `bash -c "command1; command2"` - `stdout_logfile`: Log file location for the stdout logs (`%(program_name)s` and `%(process_num)01d` can be used as variables) - `stderr_logfile` or `redirect_stderr`: Log file location for the stderr logs, or redirect all the stderr logs to the stdout stream and hence have a common file for both - `user`: The user to run the process as - `environment`: The environment variables to be set before the process is run - `numprocs`: The number of processes to run Here are some example configuration files for the GGX components: **Web Application server:** ```ini [program:corridor-app] command=INSTALL_DIR/venv-app/bin/corridor-app run stdout_logfile=/var/log/corridor/%(program_name)s.log redirect_stderr=true user=root ``` **API server:** ```ini [program:corridor-api] command= bash -c "INSTALL_DIR/venv-api/bin/corridor-api db upgrade && INSTALL_DIR/venv-api/bin/corridor-api run" stdout_logfile=/var/log/corridor/%(program_name)s.log redirect_stderr=true user=root ``` **API - Celery worker:** ```ini [program:corridor-worker-api] command=INSTALL_DIR/venv-api/bin/corridor-worker run --queue api environment= C_FORCE_ROOT=1 stdout_logfile=/var/log/corridor/%(program_name)s.log redirect_stderr=true user=root ``` **Spark - Celery worker:** ```ini [program:corridor-worker-spark] command=INSTALL_DIR/venv-api/bin/corridor-worker run --queue spark --queue quick_spark environment= C_FORCE_ROOT=1 stdout_logfile=/var/log/corridor/%(program_name)s.log redirect_stderr=true user=root ``` **Jupyter Notebook:** ```ini [program:corridor-jupyter] command=INSTALL_DIR/venv-jupyter/bin/corridor-jupyter run stdout_logfile=/var/log/corridor/%(program_name)s.log redirect_stderr=true user=root ``` ## Systemd Many Linux OS like RHEL have systemd pre-installed. To use systemd, the following steps need to be followed: - Add service file to systemd services folder. For example: `/etc/systemd/system/corridor.service` - To start service: `sudo systemctl start corridor` And to run the service on startup: `sudo systemctl enable corridor` Here are some example configuration files for the GGX components: **Web Application server:** ```ini [Unit] Description=GGX Web Application After=syslog.target network.target [Service] User=root ExecStart=/bin/bash -c 'INSTALL_DIR/venv-app/bin/corridor-app run \ >> /var/log/corridor/corridor-app.log 2>&1' Restart=always [Install] WantedBy=multi-user.target ``` **API server:** ```ini [Unit] Description=GGX API After=syslog.target network.target [Service] User=root ExecStart=/bin/bash -c 'INSTALL_DIR/venv-api/bin/corridor-api run \ >> /var/log/corridor/corridor-api.log 2>&1' Restart=always [Install] WantedBy=multi-user.target ``` **API - Celery worker:** ```ini [Unit] Description=GGX Worker API After=syslog.target network.target [Service] User=root ExecStart=/bin/bash -c 'INSTALL_DIR/venv-api/bin/corridor-worker run --queue api \ >> /var/log/corridor/corridor-worker-api.log 2>&1' Restart=always [Install] WantedBy=multi-user.target ``` **Spark - Celery worker:** ```ini [Unit] Description=GGX Worker Spark After=syslog.target network.target [Service] User=root ExecStart=/bin/bash -c 'INSTALL_DIR/venv-api/bin/corridor-worker run --queue spark --queue quick_spark \ >> /var/log/corridor/corridor-worker-spark.log 2>&1' Restart=always [Install] WantedBy=multi-user.target ``` **Jupyter Notebook:** ```ini [Unit] Description=GGX Jupyter After=syslog.target network.target [Service] User=root ExecStart=/bin/bash -c 'INSTALL_DIR/venv-api/bin/corridor-jupyter run \ >> /var/log/corridor/corridor-jupyter.log 2>&1' Restart=always [Install] WantedBy=multi-user.target ``` --- # SAML Source: https://docs.genguardx.ai/technology/self-hosting/configurations/saml/ Markdown: https://docs.genguardx.ai/technology/self-hosting/configurations/saml/index.md Description: Configure SAML single sign-on for self-hosted GGX with enterprise identity providers, authentication settings, and user access requirements. While an internal authentication is available for simple and quick installations. It is recommended to use an enterprise-grade Identity Provider (IDP) to follow the infosec requirements for your organization. GGX can integrate into IDPs and seamlessly be a tool in your organization. This section describes the use of SAML (Security Assertion Markup Language) for authentication. By using SAML, it is easy to ensure that the Platform is available only to users who are authorized to use it. It makes having a centralized Identity Provider hassle-free and ensures that the standard security practices like Single Sign-On, 2-factor Authentication, etc. are consistently applied to all the organization's applications. ## Setup To set the login based on SAML, the following information is required: From the IDP: - SSO URL: The URL endpoint to initiate Single Sign On requests Example: `http:///saml//sso` - Entity ID: The URL endpoint to fetch the SAML metadata Example: `http:///saml/` - Name ID: Unique ID to identify each user - Certificate: The X509 certificate to used to ensure any messages sent/received are trusted From GGX (Service Provider): - ACS URL: `http:///api/v1/saml/acs` - Entity ID: `http:///api/v1/saml/metadata` - Start URL: `http:///api/v1/users/saml/sso` On completing the Sign On flow on the IDP side, the information returned to GGXhould contain: - The Name ID - The following attributes: - Email (The Attribute's name can be configured with `SAML_EMAIL_ATTRIBUTE`) - List of Roles/Groups of the user (The Attribute's name can be configured with `SAML_ROLE_ATTRIBUTE`) ### Configurations In the API configurations, the following configurations need to be set: - `SAML_ENABLED = True` Needs to be set to enable SAML as the method of authentication for login. - `SAML_SETTINGS = {...}` Needs to be set as described in the configurations section to connect to the SP and IDP. The SAML_SETTINGS is a dictionary with the following information defining the SP and IDP information: ```python { # If strict is True, then the Python Toolkit will reject unsigned # or unencrypted messages if it expects them to be signed or encrypted. # Also it will reject the messages if the SAML standard is not strictly # followed. Destination, NameId, Conditions ... are validated too. "strict": true, # Enable debug mode (outputs errors). "debug": true, # Service Provider Data that we are deploying. "sp": { # Identifier of the SP entity (must be a URI) "entityId": "https:///saml/metadata/", # Specifies info about where and how the message MUST be # returned to the requester, in this case our SP. "assertionConsumerService": { # URL Location where the from the IdP will be returned "url": "https:///saml/acs", # SAML protocol binding to be used when returning the # message. "binding": "urn:oasis:names:tc:SAML:2.0:bindings:HTTP-POST" }, # Specifies info about where and how the message MUST be # returned to the requester, in this case, our SP. "singleLogoutService": { # URL Location where the from the IdP will be returned "url": "https:///saml/slo", # SAML protocol binding to be used when returning the # message. "binding": "urn:oasis:names:tc:SAML:2.0:bindings:HTTP-Redirect" }, # Specifies the constraints on the name identifier to be used to # represent the requested subject. "NameIDFormat": "urn:oasis:names:tc:SAML:2.0:nameid-format:unspecified", # The x509cert and privateKey of the SP 'x509cert': '', 'privateKey': '' }, # Identity Provider Data that we want connected with our SP. "idp": { # Identifier of the IdP entity (must be a URI) "entityId": "https:///saml/metadata", # SSO endpoint info of the IdP. (Authentication Request protocol) "singleSignOnService": { # URL Target of the IdP where the Authentication Request Message # will be sent. "url": "https:///saml/sso", # SAML protocol binding to be used when returning the # message. "binding": "urn:oasis:names:tc:SAML:2.0:bindings:HTTP-Redirect" }, # SLO endpoint info of the IdP. "singleLogoutService": { # URL Location of the IdP where SLO Request will be sent. "url": "https:///saml/sls", # SAML protocol binding to be used when returning the # message. "binding": "urn:oasis:names:tc:SAML:2.0:bindings:HTTP-Redirect" }, # Public x509 certificate of the IdP "x509cert": "" # Instead of using the whole x509cert you can use a fingerprint # (openssl x509 -noout -fingerprint -in "idp.crt" to generate it) # "certFingerprint": "" } } ``` --- # Web Server Setup Source: https://docs.genguardx.ai/technology/self-hosting/configurations/web-servers/ Markdown: https://docs.genguardx.ai/technology/self-hosting/configurations/web-servers/index.md Description: Configure Nginx, Apache, or other reverse proxies for self-hosted GGX, including routing, TLS, timeouts, compression, and application endpoints. ## Nginx Configurations This section described how to use nginx ([https://nginx.org/en/](https://nginx.org/en/)) as a web server. The Platform's server components can be made highly performant by using the lightweight Nginx as a reverse-proxy along with a WSGI server. Using nginx enabled the server to scale to a large number of users with ease. To use Nginx, it needs to be installed in the system and the daemon should be running with the appropriate site configurations setup. Here is an example nginx configuration: ```nginx server { listen 80; server_name localhost 0.0.0.0; client_max_body_size 100m; gzip on; gzip_vary on; gzip_min_length 10240; gzip_proxied expired no-cache no-store private auth; gzip_types text/plain text/css text/xml text/javascript application/x-javascript application/javascript application/xml application/json ; gzip_disable "MSIE [1-6]\."; location / { proxy_set_header X-Real-IP $remote_addr; proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; proxy_set_header X-Forwarded-Proto $scheme; proxy_set_header Host $http_host; proxy_pass http://localhost:5002; proxy_read_timeout 3600; } location /jupyter { proxy_set_header X-Real-IP $remote_addr; proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; proxy_set_header X-Forwarded-Proto $scheme; proxy_set_header Host $http_host; proxy_pass http://localhost:5003; proxy_http_version 1.1; proxy_set_header Upgrade $http_upgrade; proxy_set_header Connection "upgrade"; proxy_read_timeout 86400; } } ``` :::note If the nginx is not listening on port 80 - but listens on another port, the `proxy_set_header` for `Host` may have to be modified to `proxy_set_header Host $http_host` to ensure the correct host information is passed. ::: :::note[Note: Permission issues] In case of permission issues, ensure user permissions and SELinux is set up correctly. ::: :::note[Note: Large file uploads] If large files are expected to be uploaded, there are two settings that need to be modified
  1. add `proxy_request_buffering off`
  2. add `client_max_body_size 0;`
to server directive which expects large file size payload. This will disable nginx from buffering the large payload files and optimize disk space consumption. ::: ## Apache Configurations To setup a secure connection, update the `/etc/httpd/sites-available/corridorapp.conf` file as below: ```httpd SSLEngine On SSLCertificateFile /certs/app.crt SSLCertificateKeyFile /certs/app.key SSLCertificateChainFile /certs/ca.crt ServerName www.corridorapp.com ServerAlias corridorapp ProxyPass / http://localhost:5002/ ProxyPassReverse / http://localhost:5002/ ErrorLog /var/www/corridorapp/log/error.log CustomLog /var/www/corridorapp/log/requests.log combined ``` :::note If the httpd is not listening on port 80 (443 for SSL) - but listens on another port, the corresponding port has to be added along with `VirtualHost` keyword and the `Listen` param's value has to be appropriately updated in `/etc/httpd/conf/httpd.conf` (`/etc/httpd/conf.d/ssl.conf` for SSL). ::: :::note[Note: Permission issues] In case of permission issues, ensure user permissions and SELinux is set up correctly. ::: --- # Hardening - Security Source: https://docs.genguardx.ai/technology/self-hosting/hardening/ Markdown: https://docs.genguardx.ai/technology/self-hosting/hardening/index.md Description: Harden self-hosted GGX deployments with security controls for network access, secrets, authentication, infrastructure, data storage, and operational practices. This section describes additional setup that would be needed to make GGX secure. All the below are recommended but optional, and can be configured as needed. ## Data Storage Security GGX saves data in the following locations: - Data Lake - File Management System - Metadata Database - Jupyter Content Manager (For notebooks) Any data stored in them should be encrypted and backups should be maintained as needed. ## Network Security The following network connections are created in GGX, and should be secured: - **Web Application ↔︎ API Server**: HTTPS connection - **API Server / API - Celery ↔︎ File Management**: FTPS connection (if using FTP) - **API Server / API - Celery ↔︎ Metadata Database**: SSL connections to Database - **Spark - Celery / Jupyter ↔︎ Spark**: Kerberos - **GGX package ↔︎ API Server**: HTTPS connection - **Jupyter ↔︎ Web Application**: HTTPS connection - **End User ↔︎ Web Application / Jupyter**: HTTPS / WSS connection ## Securing each component This section describes the steps to follow for each of the GGX components to ensure it can be accessed securely. ### Common - Ensure that all configuration and installation files are readable only by the user that the process is running with ### Web Application Server - Ensure that the application is served using a standard web server like Nginx or Apache httpd in front of the WSGI server - Setup the secure HTTPS protocol at the WSGI Server or the Web Server using an SSL certificate - When setting up HTTPS, also set the JWT_COOKIE_SECURE configuration to ensure JWT cookies are sent in a secure manner - Set a strong and unique SECRET_KEY - It should use a reliable Authentication Provider (Avoid using the inbuilt authentication provider) ### API Server - Ensure that the application is served using a standard web server like Nginx or Apache httpd in front of the WSGI server - Setup the secure HTTPS protocol at the WSGI Server or the Web Server using an SSL certificate - Set a strong and unique SECRET_KEY - Ensure API Keys are set to ensure only authorized access to the APIs ### API - Celery worker No specific steps are required for the API - Celery workers as no other component connects to it directly. ### Spark - Celery worker - Ensure that the cluster is Kerberized - Ensure the standard security practices for Spark are followed as described in the [Spark - Security](https://spark.apache.org/docs/latest/security.html) documentation. ### Jupyter Notebook - Ensure that the application is served using a standard web server like Nginx or Apache httpd in front of the WSGI server - Setup the secure HTTPS protocol at the WSGI Server or the Web Server using an SSL certificate - Ensure the standard security practices for Jupyter are followed as described in the [Jupyter - Security](https://jupyter.org/security) documentation. ### File Management For Local File System: Ensure that the Hard disk being used is encrypted. For FTP: Ensure the FTPS protocol is being used and the underlying data is encrypted ### Metadata Database (SQL RDBMS) - Ensure that the connection to the SQL database is secured using any of the authentication methods available to the RDBMS. - Ensure that the Database is encrypted. - Ensure the standard security practices for the RDBS are followed as described in its documentation. ### Authentication Provider - For LDAP: Ensure the secure LDAPS is used to create connections - For SAML: Ensure a valid x509 certificate is used to authenticate messages being sent/received --- # AWS Source: https://docs.genguardx.ai/technology/self-hosting/installation/aws/ Markdown: https://docs.genguardx.ai/technology/self-hosting/installation/aws/index.md Description: Deploy self-hosted GGX on AWS using ECS Fargate, EFS, load balancers, networking, IAM, database services, and container registry configuration. Use this page to choose and configure an AWS deployment path for GGX. ## Recommended AWS Paths | Path | Use when | Primary docs | |---|---|---| | EKS | You already operate Kubernetes or need namespace isolation and Kubernetes-native operations | [Kubernetes](../kubernetes/) | | ECS Fargate | You want AWS-managed containers without managing Kubernetes nodes | [Terraform](../terraform/) | | EC2 or other VMs | You want a traditional VM-based install and direct OS control | [Manual](../manual/) | ## EKS-Specific Configuration EKS uses the shared [GGX Kubernetes manifests](https://github.com/corridor/kubernetes-ggx). Start with the [Kubernetes](../kubernetes/) page, then apply the AWS-specific requirements below. ### Required AWS Services - **Amazon EKS** for the managed Kubernetes cluster. - **Amazon RDS for PostgreSQL** for GGX metadata. - **Amazon EFS** for read-write-many persistent volumes. - **Application Load Balancer** through AWS Load Balancer Controller. - **Amazon VPC** with private subnets for workloads and controlled public ingress. - **IAM** for cluster roles, controller permissions, and workload identities. Optional but common services: - **Route 53** for DNS. - **AWS Certificate Manager** for TLS certificates. - **Secrets Manager** for sensitive configuration. - **CloudWatch** for logs, metrics, and alarms. - **AWS WAF** for edge protection. ### Permissions The deploying role or CI identity needs permission to manage: - EKS clusters and managed node groups. - VPCs, subnets, route tables, NAT Gateways, and security groups. - IAM roles, policies, and IAM Roles for Service Accounts (IRSA). - EFS file systems, mount targets, and access points. - RDS instances, subnet groups, and security groups. - ALB listeners, target groups, and ingress-related resources. - Route 53 records and ACM certificates when DNS and TLS are managed in AWS. ### Cluster Add-ons Install or enable these before applying the GGX overlay: - AWS Load Balancer Controller for ALB-backed ingress. - EFS CSI Driver for persistent volumes. - cert-manager if TLS is issued by Kubernetes. - Cluster Autoscaler or Karpenter for node scaling. - CloudWatch Container Insights or another approved observability stack. ### Networking Production EKS deployments should normally use private worker nodes with outbound internet access through NAT. Security groups must allow: - ALB to reach the GGX app and Jupyter services. - GGX pods to reach RDS PostgreSQL. - GGX pods to reach EFS mount targets. - Pods to pull GGX images from the configured registry. ## ECS Fargate With Terraform The [GGX AWS Terraform module](https://github.com/corridor/terraform-aws-ggx) deploys GGX on ECS Fargate. This is the main non-Kubernetes AWS path. The Fargate deployment uses: - A single ECS service with a task definition containing `corridor-migration`, `corridor-app`, `corridor-worker`, and `corridor-jupyter`. - Application Load Balancer routing `/` to `corridor-app` on port `5002`. - Application Load Balancer routing `/jupyter` to `corridor-jupyter` on port `5003`. - EFS for shared persistent storage. - CloudWatch logs. - IAM task execution and task roles. Configure the module with the GGX image, hostname, ACM certificate ARN, database URL, and license key. Then run: ```bash terraform init terraform plan terraform apply ``` Useful ECS operations: ```bash aws logs tail /ecs/corridor --follow aws ecs update-service --cluster corridor --service corridor --force-new-deployment aws ecs describe-services --cluster corridor --services corridor ``` ## EC2 Or VM-Based Installs An EC2 deployment follows the [Manual](../manual/) path. The EC2 installation pattern is: 1. Launch an EC2 instance sized from the [minimum requirements](../minimum-requirements/), commonly `t3.2xlarge` or larger for all-in-one deployments. 2. Create an RDS PostgreSQL database. 3. Install system dependencies such as Python 3.11, Java 8 for Spark, Nginx, and unzip. 4. Extract the GGX installation bundle. 5. Install the `app`, `api`, `worker-api`, `worker-spark`, and `jupyter` components. 6. Configure `/opt/corridor/instances/default/config/api_config.py`. 7. Run `corridor-api db upgrade`. 8. Create systemd services and start the components. Use EC2 when you need direct host access or your organization standardizes on VM operations. Use EKS or ECS Fargate when you want managed container operations. ## Security Notes - Do not deploy with the AWS account root user. - Store application secrets in Secrets Manager or an approved secret store. - Use private subnets for application workloads and databases. - Enable encryption at rest for RDS and EFS. - Use least-privilege IAM roles for controllers, tasks, and operations. - Enable CloudWatch logs and billing alerts before production rollout. --- # Azure Source: https://docs.genguardx.ai/technology/self-hosting/installation/azure/ Markdown: https://docs.genguardx.ai/technology/self-hosting/installation/azure/index.md Description: Deploy self-hosted GGX on Azure using Container Apps, Azure Files, managed networking, database services, container registry, and Terraform-based infrastructure. Use this page to choose and configure an Azure deployment path for GGX. ## Recommended Azure Paths | Path | Use when | Primary docs | |---|---|---| | AKS | You already operate Kubernetes or need Kubernetes-native scaling and operations | [Kubernetes](../kubernetes/) | | Azure Container Apps | You want Azure-managed containers without managing a Kubernetes cluster | [Terraform](../terraform/) | | Azure VMs | You want a traditional VM-based install and direct OS control | [Manual](../manual/) | ## AKS-Specific Configuration AKS uses the shared [GGX Kubernetes manifests](https://github.com/corridor/kubernetes-ggx). Start with the [Kubernetes](../kubernetes/) page, then apply the Azure-specific requirements below. ### Required Azure Services - **Azure Kubernetes Service** for the managed Kubernetes cluster. - **Azure Database for PostgreSQL** for GGX metadata. - **Azure Files Premium** or another approved read-write-many storage provider. - **Azure Virtual Network** for private networking. - **Azure DNS** or another DNS provider. Optional but common services: - **Azure Key Vault** for secrets. - **Azure Monitor** for logs and metrics. - **Azure Front Door** or Web Application Firewall for edge protection. - **Application Gateway Ingress Controller** when your platform standardizes on Application Gateway. ### Permissions The deploying identity needs permission to manage: - AKS clusters and node pools. - Virtual networks, subnets, route tables, private DNS zones, and network security groups. - Managed identities and role assignments. - Azure Files storage accounts and file shares. - PostgreSQL servers, firewall rules, and private endpoints when used. - DNS records and TLS certificate resources when managed in Azure. - Key Vault secrets when application secrets are stored there. ### Cluster Add-ons Install or enable these before applying the GGX overlay: - Azure Files CSI Driver. - NGINX Ingress Controller or Application Gateway Ingress Controller. - cert-manager if TLS is issued from the cluster. - Azure Monitor Container Insights or another approved observability stack. - Network Policy if your environment requires pod-to-pod controls. ### Networking Production AKS deployments should normally use controlled ingress and private connectivity to PostgreSQL and storage. Network security groups and database firewall rules must allow: - Ingress controller to reach `corridor-app` and `corridor-jupyter`. - GGX pods to reach Azure Database for PostgreSQL. - GGX pods to mount Azure Files. - Pods to pull GGX images from the configured registry. ## Azure Container Apps With Terraform The [GGX Azure Terraform module](https://github.com/corridor/terraform-azurerm-ggx) deploys GGX on Azure Container Apps. This is the main non-Kubernetes Azure container path. The module provisions or configures: - Container Apps for the GGX app, worker, Jupyter, PostgreSQL-facing configuration, and Nginx routing. - Azure Files for shared state. - Optional dedicated workload profiles when higher memory or predictable capacity is required. - Outputs for the app URL, Jupyter URL, Container App Environment, storage account, and database details. Important inputs include the Azure region, ACR login server, ACR service principal credentials, image name, image version, GGX license key, database admin password, and optional workload profile. ```bash terraform init terraform plan terraform apply ``` ## Azure VM-Based Installs An Azure VM deployment follows the [Manual](../manual/) path. The Azure VM installation pattern is: 1. Create a resource group and Azure VM, commonly `Standard_D8s_v3` or larger for an all-in-one deployment. 2. Attach and mount a data disk for `/opt/corridor` and application state. 3. Create Azure Database for PostgreSQL. 4. Install Python 3.11, Java 8 for Spark, Nginx, and unzip. 5. Extract the GGX installation bundle. 6. Install the `app`, `api`, `worker-api`, `worker-spark`, and `jupyter` components. 7. Configure database and application settings. 8. Run database migrations. 9. Create systemd services and start the components. Use Azure VMs when you need direct host access or your organization standardizes on VM operations. Use AKS or Azure Container Apps when you want managed container operations. ## Security Notes - Use managed identities where possible. - Store secrets in Key Vault or an approved secret store. - Use private networking for PostgreSQL and storage. - Enable encryption at rest for database and file storage. - Restrict SSH access and use just-in-time access where available. - Enable Azure Monitor and alerting before production rollout. --- # Docker-based Source: https://docs.genguardx.ai/technology/self-hosting/installation/docker-based/ Markdown: https://docs.genguardx.ai/technology/self-hosting/installation/docker-based/index.md Description: Run self-hosted GGX with Docker-based deployments for local, staging, or controlled environments using containers, volumes, configuration, and service orchestration. Use a Docker-based deployment when you already operate Docker hosts or Docker Compose and want a containerized GGX installation without adopting Kubernetes. GGX provides production-ready Dockerfile templates with the installation bundle. Because Docker networking, storage, reverse proxies, and secret management vary significantly by organization, work with GGX support to align the final compose or runtime configuration with your environment. ## Components | Component | Implementation options | |---|---| | Web application | GGX app container behind a reverse proxy | | API and workers | GGX worker containers for API and Spark queues | | Jupyter | Jupyter container or integration with an existing notebook service | | File management | Host mount, Docker volume, or network-attached persistent storage | | Metadata database | External PostgreSQL, Oracle, SQL Server, or managed database service | | SSL certificates | Mounted read-only from host storage or injected during image build | | Platform configuration | Mounted read-only files, environment variables, or approved secret store | ## Requirements - Docker Engine 20.10 or later. - Docker Compose or an equivalent orchestration process if running multiple containers manually. - Persistent storage for uploads, notebooks, data, state, and backups. - Metadata database that meets the [minimum requirements](../minimum-requirements/). - Reverse proxy such as Nginx, Apache, an ingress appliance, or a cloud load balancer. - TLS certificate for browser-facing traffic. ## Deployment Flow 1. Prepare the host, Docker runtime, persistent storage, database, DNS, and TLS certificate. 2. Build or pull the GGX images provided for your release. 3. Configure shared volumes for data, notebooks, uploads, Jupyter state, and backups. 4. Configure application settings through mounted config files or environment variables. 5. Start the app, worker, Jupyter, and supporting containers. 6. Run database migrations. 7. Configure the reverse proxy so `/` reaches the app service and `/jupyter` reaches the Jupyter service. 8. Verify logs, health checks, persistent storage, and user login. ## Configuration Notes - Keep the database outside the app container for production deployments. - Use named volumes or network storage rather than ephemeral container filesystems. - Store secrets in your platform secret store instead of hard-coding them in compose files. - Separate app, worker, and Jupyter logs so operations teams can troubleshoot independently. - Align CPU and memory reservations with the [minimum requirements](../minimum-requirements/). ## When To Choose Another Path - Use [Kubernetes](../kubernetes/) for managed container orchestration on AKS, GKE, or EKS. - Use [Terraform](../terraform/) for cloud-managed container services such as ECS Fargate, Azure Container Apps, or Cloud Run. - Use [Manual](../manual/) for VM or bare-metal installations that do not use containers. --- # GCP Source: https://docs.genguardx.ai/technology/self-hosting/installation/gcp/ Markdown: https://docs.genguardx.ai/technology/self-hosting/installation/gcp/index.md Description: Deploy self-hosted GGX on Google Cloud using Cloud Run, GKE, Cloud SQL, Cloud Storage, networking, service accounts, and Terraform modules. Use this page to choose and configure a Google Cloud deployment path for GGX. ## Recommended GCP Paths | Path | Use when | Primary docs | |---|---|---| | GKE | You already operate Kubernetes or need Kubernetes-native scaling and operations | [Kubernetes](../kubernetes/) | | Cloud Run | You want Google-managed containers without managing a Kubernetes cluster | [Terraform](../terraform/) | | Compute Engine VMs | You want a traditional VM-based install and direct OS control | [Manual](../manual/) | ## GKE-Specific Configuration GKE uses the shared [GGX Kubernetes manifests](https://github.com/corridor/kubernetes-ggx). Start with the [Kubernetes](../kubernetes/) page, then apply the GCP-specific requirements below. ### Required GCP Services - **Google Kubernetes Engine** for the managed Kubernetes cluster. - **Cloud SQL for PostgreSQL** for GGX metadata. - **Filestore** or another approved read-write-many storage provider. - **VPC** with private networking for cluster, database, and storage connectivity. - **Cloud NAT** when private nodes need outbound internet access. Optional but common services: - **Cloud DNS** for DNS. - **Secret Manager** for sensitive configuration. - **Cloud Monitoring and Cloud Logging** for observability. - **Cloud Armor** for edge protection. - **Cloud CDN** for static asset caching. ### Permissions The deploying identity needs permission to manage: - GKE clusters and node pools. - VPC networks, subnets, firewall rules, routers, and Cloud NAT. - Service accounts and IAM bindings. - Cloud SQL instances, databases, and users. - Filestore instances or the selected storage provider. - Cloud DNS records and certificate resources when managed in GCP. - Secret Manager secrets when application secrets are stored there. Enable the required APIs before deployment, including Kubernetes Engine, Compute Engine, Cloud SQL, and any storage, DNS, certificate, or monitoring APIs used by your environment. ### Cluster Add-ons Install or enable these before applying the GGX overlay: - Ingress controller appropriate for your load balancer strategy. - cert-manager if TLS is issued from the cluster. - NFS or Filestore CSI/provisioning components for read-write-many volumes. - Workload Identity if pods need direct access to Google Cloud services. - Cloud Monitoring integration or another approved observability stack. ### Networking Production GKE deployments should normally use private nodes with outbound internet through Cloud NAT. Firewall rules must allow: - Ingress load balancer to reach `corridor-app` and `corridor-jupyter`. - GGX pods to reach Cloud SQL. - GGX pods to mount Filestore or the selected storage provider. - Pods to pull GGX images from the configured registry. ## Cloud Run With Terraform The [GGX Google Cloud Terraform module](https://github.com/corridor/terraform-google-ggx) deploys GGX on Cloud Run. This is the main non-Kubernetes GCP container path. The module provisions or configures: - `corridor-migration` as a Cloud Run Job. - `corridor-app` as a public Cloud Run service. - `corridor-worker` as an internal Cloud Run service with minimum instances. - `corridor-jupyter` as a public Cloud Run service. - Cloud SQL for PostgreSQL. - Cloud Storage for shared file-backed state. - Direct VPC egress from Cloud Run to private services. - External HTTPS load balancer with serverless NEGs so `/` routes to the app and `/jupyter` routes to Jupyter. Configure the module with the project ID, image, hostname, database password, license key, and SMTP values if email is required. ```bash terraform init terraform plan terraform apply ``` After apply, point DNS at the reserved load balancer IP and wait for the managed certificate to become active. ## Compute Engine VM-Based Installs A Compute Engine deployment follows the [Manual](../manual/) path. The Compute Engine installation pattern is: 1. Create a Compute Engine VM, commonly `n2-standard-8` or similar for all-in-one deployments. 2. Create Cloud SQL for PostgreSQL. 3. Install Python 3.11, Java 8 for Spark, Nginx, and unzip. 4. Extract the GGX installation bundle. 5. Install the `app`, `api`, `worker-api`, `worker-spark`, and `jupyter` components. 6. Configure database and application settings. 7. Run database migrations. 8. Create systemd services and start the components. Use Compute Engine when you need direct host access or your organization standardizes on VM operations. Use GKE or Cloud Run when you want managed container operations. ## Security Notes - Use service accounts with least-privilege IAM roles. - Store secrets in Secret Manager or an approved secret store. - Use private IP connectivity for Cloud SQL where possible. - Enable encryption, automated backups, and monitoring. - Restrict SSH access and prefer OS Login or IAP where available. - Enable Cloud Logging and alerting before production rollout. --- # Kubernetes Source: https://docs.genguardx.ai/technology/self-hosting/installation/kubernetes/ Markdown: https://docs.genguardx.ai/technology/self-hosting/installation/kubernetes/index.md Description: Deploy self-hosted GGX on Kubernetes using Kustomize manifests, namespaces, persistent volumes, ingress, secrets, and provider-specific cluster settings. Use Kubernetes when you want a cloud-native GGX deployment with managed rollout, namespace isolation, persistent volumes, and standard cluster operations. GGX provides cloud-agnostic [GGX Kubernetes manifests](https://github.com/corridor/kubernetes-ggx). The manifests use Kustomize and can run on managed Kubernetes services including Azure Kubernetes Service (AKS), Google Kubernetes Engine (GKE), and Amazon Elastic Kubernetes Service (EKS). ## Deployment Shape The Kubernetes deployment keeps the same GGX service split used by the other installation paths: ```text Kubernetes cluster ├── ggx namespace │ ├── corridor-app │ ├── corridor-worker │ └── corridor-jupyter ├── persistent volumes for data, uploads, notebooks, state, and backups └── ingress routing / to corridor-app and /jupyter to corridor-jupyter ``` The `kubernetes-ggx` repository contains reusable manifests in `base/` and a deployable example overlay in `overlays/example/`. Create environment-specific overlays such as `overlays/prod`, `overlays/staging`, or separate team overlays when you need isolated deployments. ## Prerequisites - A Kubernetes cluster with enough CPU, memory, and persistent storage for the [minimum requirements](../minimum-requirements/). - `kubectl` access with permission to create namespaces, secrets, config maps, deployments, services, ingress resources, and persistent volume claims. - GGX container registry credentials from GGX support. - PostgreSQL metadata database connectivity. - Persistent storage that supports the access pattern used by your deployment. - DNS name and TLS certificate strategy for the public application endpoint. ## Quickstart Create the namespace before creating namespace-scoped objects: ```bash kubectl create namespace ggx ``` Create the image pull secret using the registry credential JSON provided by GGX: ```bash kubectl create secret docker-registry corridor-registry-secret \ --docker-server=us-central1-docker.pkg.dev \ --docker-username=_json_key \ --docker-password="$(cat /tmp/corridor-registry-key.json)" \ --namespace ggx ``` Apply an overlay after you have reviewed and customized it: ```bash kubectl apply -k overlays/example ``` Verify the rollout: ```bash kubectl get pods -n ggx kubectl get svc -n ggx kubectl get ingress -n ggx ``` ## Overlay Configuration Before applying the overlay, configure the deployment for your environment: - Set the GGX image tag in `overlays/example/kustomization.yaml`. - Set the public hostname in `overlays/example/kustomization.yaml`. - Set database, authentication, and application settings in `overlays/example/configs/api_config.py`. - Update persistent volume claim patches if your cluster uses a different read-write-many storage class. - Configure TLS secrets, ingress annotations, and timeout or gzip settings in the ingress manifest. - Tune CPU and memory requests and limits in the service manifests. ## Provider Notes ### AKS AKS deployments usually need: - Azure RBAC permissions for AKS, virtual networks, managed identities, Azure Files, Azure Database for PostgreSQL, DNS, Key Vault, and Azure Monitor. - Azure CNI or another networking mode approved by your platform team. - Azure Files CSI Driver or an equivalent read-write-many storage provider. - NGINX Ingress Controller or Application Gateway Ingress Controller. - cert-manager or an Azure-managed certificate process. - Network security group and database firewall rules allowing the cluster to reach PostgreSQL and shared storage. Use the [Azure](../azure/) page for AKS-specific service choices and Terraform alternatives. ### GKE GKE deployments usually need: - IAM permissions for GKE, Compute Engine networking, Cloud SQL, Filestore or another RWX storage service, Secret Manager, Cloud DNS, and Cloud Monitoring. - Required APIs enabled, including Kubernetes Engine, Compute Engine, Cloud SQL, and any storage or DNS APIs used by the deployment. - Private cluster egress through Cloud NAT when worker nodes do not have public IPs. - Filestore, another NFS provider, or a compatible RWX storage class for persistent volumes. - Workload Identity if pods need direct access to Google Cloud services. - Ingress and certificate configuration appropriate for your load balancer choice. Use the [GCP](../gcp/) page for GKE-specific service choices and Terraform alternatives. ### EKS EKS deployments usually need: - IAM permissions for EKS, IAM, VPC, EC2, EFS, RDS, Elastic Load Balancing, ACM, Route 53, Secrets Manager, and CloudWatch. - AWS Load Balancer Controller for ALB-backed ingress. - EFS CSI Driver for read-write-many persistent volumes. - IAM Roles for Service Accounts (IRSA) for controllers and any workload permissions. - VPC CNI, Cilium, or another approved pod networking configuration. - Security groups allowing cluster workloads to reach RDS, EFS mount targets, and any external dependencies. Use the [AWS](../aws/) page for EKS-specific service choices and Terraform alternatives. ## Recommended Starting Cluster | Provider | Cluster type | Starting nodes | Autoscaling | Node size | Disk | Notes | |---|---|---:|---|---|---:|---| | AKS | Standard | 1 | min 1, max 3 | `Standard_D8s_v3` | 100 GB | Use Azure CNI and approved virtual network | | GKE | Standard | 1 | min 1, max 3 | `e2-standard-8` | 100 GB | Enable IP aliasing and approved VPC/subnet | | EKS | Standard | 1 | min 1, max 3 | `m5.xlarge` or similar | 100 GB | Attach to approved VPC/subnet | A single node is usually enough for proof-of-concept usage. Production deployments should size node pools, storage, database, and autoscaling limits based on expected user count, data volume, and background job concurrency. ## Operations Common Kubernetes operations: ```bash kubectl logs -n ggx deploy/corridor-app kubectl logs -n ggx deploy/corridor-worker kubectl exec -it -n ggx deploy/corridor-app -- /bin/bash kubectl rollout restart deployment corridor-app -n ggx kubectl rollout status deployment corridor-app -n ggx kubectl get pvc -n ggx ``` If pods show `ImagePullBackOff`, check that `corridor-registry-secret` exists in the application namespace and contains the current registry credentials. --- # Manual Source: https://docs.genguardx.ai/technology/self-hosting/installation/manual/ Markdown: https://docs.genguardx.ai/technology/self-hosting/installation/manual/index.md Description: Install self-hosted GGX manually on VMs, bare metal, or cloud instances with direct control over services, databases, storage, web servers, and process managers. Use a manual installation when you are deploying GGX directly on VMs, bare metal, or cloud instances and want direct control over the operating system, process manager, web server, database, and storage. This path applies to on-premises servers and VM-based cloud deployments such as EC2, Azure VMs, and Compute Engine. ## Prerequisites Before starting, ensure the [minimum requirements](../minimum-requirements/) are met. You also need: - GGX installation bundle. - Linux host access with privilege to install packages and create services. - Python 3.11 or later. - Java 8 or later for Spark worker functionality. - Metadata database: PostgreSQL 11.7 or later, Oracle 19 or later, or SQL Server 2016 or later. - Persistent file storage. - Nginx, Apache, or another production reverse proxy. - TLS certificate for browser-facing traffic. ## Components The installation bundle provides command-line entry points for each component: - Web application server: `corridor-app` - API server: `corridor-api` - API Celery worker: `corridor-worker` - Spark Celery worker: `corridor-worker` - Jupyter Notebook server: `corridor-jupyter` ## Install Components Extract the bundle and run the installer for each required component: ```sh unzip corridor-bundle.zip sudo ./corridor-bundle/install [app | api | worker-api | worker-spark | jupyter] ``` The installer creates: - A component-specific virtual environment under the installation path. - Configuration files under the selected instance name. - Component entry points for running services. Check installer options with: ```text usage: install [-h] [-i INSTALL_DIR] [-n NAME] component positional arguments: component The component to install. Possible values are: api, app, worker-api, worker-spark, jupyter optional arguments: -h, --help show this help message and exit -e EXTRAS [EXTRAS ...], --extras EXTRAS [EXTRAS ...] The extra packages to install -i INSTALL_DIR, --install-dir INSTALL_DIR The location to install the GGX package. Default value: /opt/corridor --overwrite Whether to overwrite the configs if already present. Default behavior is to create config files only if they don't already exist. ``` ## Configure The API Update the API configuration, usually at: ```text INSTALL_DIR/instances/INSTANCE_NAME/config/api_config.py ``` Configure the database connection string, file storage, authentication, email, and other platform settings required by your environment. Run database migrations: ```sh INSTALL_DIR/venv-api/bin/corridor-api db upgrade ``` ## Run Components ### Web Application Server - Configuration file: `INSTALL_DIR/instances/INSTANCE_NAME/config/app_config.py` - Run command: `INSTALL_DIR/venv-app/bin/corridor-app run` - WSGI application: `corridor_app.wsgi:app` Set `WSGI_SERVER` to `gunicorn` or `auto` for production. Do not use Werkzeug for production. ### API Server - Configuration file: `INSTALL_DIR/instances/INSTANCE_NAME/config/api_config.py` - Run command: `INSTALL_DIR/venv-api/bin/corridor-api run` - WSGI application: `corridor_api.wsgi:app` Set `WSGI_SERVER` to `gunicorn` or `auto` for production. Do not use Werkzeug for production. ### API Worker - Configuration file: `INSTALL_DIR/instances/INSTANCE_NAME/config/api_config.py` - Run command: `INSTALL_DIR/venv-api/bin/corridor-worker run --queue api` If the process runs as root, set `C_FORCE_ROOT=1`. ### Spark Worker - Configuration file: `INSTALL_DIR/instances/INSTANCE_NAME/config/api_config.py` - Run command: `INSTALL_DIR/venv-api/bin/corridor-worker run --queue spark --queue quick_spark` The Spark worker should run on a machine configured as a Spark gateway or edge node. It should be able to import `pyspark` and reach the target Spark cluster. ### Jupyter Notebook - Jupyter Hub configuration file: `INSTALL_DIR/instances/INSTANCE_NAME/config/jupyterhub_config.py` - Jupyter Notebook configuration file: `INSTALL_DIR/instances/INSTANCE_NAME/config/jupyter_notebook_config.py` - Run command: `INSTALL_DIR/venv-jupyter/bin/corridor-jupyter run` ## Process Management Use systemd, Supervisor, or your standard process manager. A typical systemd service follows this shape: ```ini [Unit] Description=GGX API Server After=network.target [Service] Type=simple User=corridor Group=corridor Environment=CORRIDOR_CONFIG_DIR=/opt/corridor/instances/default/config Environment=WSGI_SERVER=gunicorn ExecStart=/opt/corridor/venv-api/bin/corridor-api run Restart=always [Install] WantedBy=multi-user.target ``` After creating services: ```sh sudo systemctl daemon-reload sudo systemctl enable corridor-app corridor-api corridor-worker-api corridor-worker-spark corridor-jupyter sudo systemctl start corridor-app corridor-api corridor-worker-api corridor-worker-spark corridor-jupyter ``` ## Reverse Proxy Place Nginx, Apache, or an approved load balancer in front of the application services: - Route browser traffic for the main app to the web application server. - Route API traffic to the API server if your topology separates app and API. - Route `/jupyter` to the Jupyter service. - Terminate TLS at the reverse proxy or load balancer. - Set secure cookie and secret configuration values before production use. ## Operations ```sh sudo systemctl status corridor-api sudo systemctl restart corridor-api journalctl -u corridor-api -f ``` For cloud VM deployments, use the provider page for required cloud services: - [AWS](../aws/) for EC2, RDS, EBS or EFS, Route 53, and AWS security controls. - [Azure](../azure/) for Azure VMs, Azure Database for PostgreSQL, Azure Files, and Azure networking. - [GCP](../gcp/) for Compute Engine, Cloud SQL, Filestore or persistent disks, and GCP networking. --- # Minimum Requirements Source: https://docs.genguardx.ai/technology/self-hosting/installation/minimum-requirements/ Markdown: https://docs.genguardx.ai/technology/self-hosting/installation/minimum-requirements/index.md Description: Review the minimum CPU, memory, storage, database, Spark, Python, Java, web server, and process management requirements for self-hosted GGX components. This section describes the minimum requirements that are needed for a GGX installation. Broadly, the components involved are: - Web Application & Worker - Spark Worker - Jupyter Notebook - File Management - Metadata Database (SQL RDBMS) For very simple installations, all of these could be installed on the same machine, we recommend keeping them separate to simplify scalability needs. ## Web Application A flask application which serves the User Interface and Web APIs which are accessible to users via the browser. It also includes a worker process for long running tasks in the API. This component has 2 processes: `corridor-app` and `corridor-worker` ### Requirements - RAM: 4 GB - Processor: 4 CPU - Installation storage space: 20 GB - Python 3.11+ Optional: - Web Server - Example: Nginx - Process Management - Example: Supervisor or Systemd ## Spark Worker Worker to handle any jobs triggered by users which are asynchronously. It is recommended to have at least 2 workers and increase concurrency as required. :::note This needs to be installed on a machine that is configured as a Spark Gateway (i.e. A master node or an edge node of the cluster). This is not the data nodes of the cluster itself. The worker process should be able to import the `pyspark` module. ::: ### Requirements - RAM: 16 GB - Processor: 8 CPU - HDFS storage space: 500 GB (depends on the data being processed, HDFS space to handle shuffles need to be considered too) - Python 3.11+ - Java 8+ - Spark 3.3+ Optional: - Process Management - Example: Supervisor or Systemd ## Jupyter Notebook A notebook for free-form analytical usage. We provide Jupyter Notebooks out-of-the-box but can integrate with existing notebook solutions too. :::note This needs to be installed on a machine that is configured as a Spark Gateway (i.e. A master node or an edge node of the cluster). This is not the data nodes of the cluster itself. The jupyter notebook kernel should be able to import the `pyspark` module. ::: ### Requirements - RAM: 4 GB for base services and more as per usage by users - Processor: 4 CPU and more as per usage by users - Installation storage space: 10 GB - Python 3.11+ - Spark 3.3+ Optional: - Process Management - Example: Supervisor or Systemd ## File Management A file system management to store and retrieve files. A NAS storage that can be mounted on all servers and be accessible by all services is ideal. ### Requirements - File storage space: 50 GB ## Metadata Database This serves as an internal RDBMS to store the state of the application and various user information. ### Requirements - RAM: 2 GB - Processor: 2 CPU - Database storage space: 5 GB - SQL Databases supported: - Oracle 19+ - MSSQL 2016+ - Postgres 11.7+ --- # Terraform Source: https://docs.genguardx.ai/technology/self-hosting/installation/terraform/ Markdown: https://docs.genguardx.ai/technology/self-hosting/installation/terraform/index.md Description: Provision self-hosted GGX infrastructure as code with Terraform modules for AWS, Azure, and Google Cloud managed container deployments. Use Terraform when you want GGX infrastructure provisioned as code on a supported cloud. GGX maintains cloud-specific Terraform repositories for managed container deployments: - [GGX AWS Terraform module](https://github.com/corridor/terraform-aws-ggx) for AWS ECS Fargate. - [GGX Azure Terraform module](https://github.com/corridor/terraform-azurerm-ggx) for Azure Container Apps. - [GGX Google Cloud Terraform module](https://github.com/corridor/terraform-google-ggx) for Google Cloud Run. These modules are separate from the [Kubernetes](../kubernetes/) manifests. Use Kubernetes for AKS, GKE, or EKS clusters. Use Terraform when you want cloud-managed container services and the surrounding cloud infrastructure created through IaC. ## Common Workflow Each repository follows the same Terraform workflow: ```bash cp terraform.tfvars.example terraform.tfvars # Edit terraform.tfvars with cloud, image, database, hostname, and license values. terraform init terraform plan terraform apply ``` Keep `terraform.tfvars` and state files out of source control unless your organization has an approved secrets and backend workflow. For production, configure a remote backend such as S3, Azure Storage, or GCS and restrict state access because state may contain sensitive values. ## AWS Module The AWS module runs GGX on ECS Fargate. It provisions or configures: - One ECS service on an ECS cluster. - A Fargate task definition with `corridor-migration`, `corridor-app`, `corridor-worker`, and `corridor-jupyter`. - Application Load Balancer routing `/` to the app container and `/jupyter` to Jupyter. - EFS file system, mount targets, and access points for shared persistent state. - IAM task execution and task roles. - CloudWatch log group. - Security groups for ALB, ECS tasks, and EFS. Primary configuration values include: - `image` - `hostname` - `certificate_arn` - `database_url` - `license_key` See the [AWS](../aws/) page for AWS service and permission guidance. ## Azure Module The Azure module deploys GGX on Azure Container Apps. It provisions or configures: - Container Apps for the app, worker, Jupyter, PostgreSQL-facing configuration, and Nginx routing. - Azure Files for shared storage. - Optional dedicated workload profiles when the default consumption profile is not enough. - Resource group, Container App Environment, storage account, and database-related outputs. Primary configuration values include: - `resource_group_name` - `location` - `acr_login_server` - `acr_sp_client_id` - `acr_sp_client_secret` - `image_name` - `image_version` - `corridor_license_key` - `db_admin_password` - `app_workload_profile` See the [Azure](../azure/) page for Azure service and permission guidance. ## Google Cloud Module The Google Cloud module runs GGX on Cloud Run and maps the Kubernetes application shape to managed Google Cloud services. It provisions or configures: - `corridor-migration` as a Cloud Run Job. - `corridor-app`, `corridor-worker`, and `corridor-jupyter` as Cloud Run services. - Cloud SQL for PostgreSQL. - Cloud Storage for shared file-backed state. - Direct VPC egress for private service connectivity. - External HTTPS load balancer with serverless NEGs so `/` routes to the app and `/jupyter` routes to Jupyter. - Service account and IAM bindings. Primary configuration values include: - `project_id` - `image` - `hostname` - `db_password` - `license_key` - SMTP values when email notifications are required See the [GCP](../gcp/) page for Google Cloud service and permission guidance. ## When To Use Terraform | Requirement | Recommended path | |---|---| | Managed Kubernetes on AKS, GKE, or EKS | [Kubernetes](../kubernetes/) | | AWS serverless containers | Terraform AWS ECS Fargate module | | Azure managed containers | Terraform Azure Container Apps module | | Google Cloud managed containers | Terraform Cloud Run module | | Existing VMs or bare metal | [Manual](../manual/) | | Existing Docker host or compose environment | [Docker-based](../docker-based/) | --- # Backups & Restore Source: https://docs.genguardx.ai/technology/self-hosting/scaling/backups/ Markdown: https://docs.genguardx.ai/technology/self-hosting/scaling/backups/index.md Description: Back up and restore self-hosted GGX metadata databases, file management storage, data lake outputs, settings, and component configuration files. To perform backups, it is important to understand the data that the platform uses/saves. Data that the platform writes and the systems used for persistent storage are: - Metadata Database - File Management - Data Lake - Settings ## Backing up the systems ### Metadata Database The entire database used for GGX needs to be backed up. There are 2 ways to run the backup: - Use the standard backup manager as recommended/preferred for your RDBMS. For example: `expdp` for Oracle. - Use the `corridor-API db export` command in GGX to save your database into an SQLite file The `corridor-api db export` command will create an SQLite file which is a copy of your database. It includes various referential guarantees that databases provide like unique constraints, foreign keys, etc. It can be directly used as an embedded database with GGX or can be used to import back into your RDBMS of choice. This also supports converting from 1 RBMS system to another. ### File Management Depending on the system that is being used, backups need to be created. If FTP or NFS is being used - the appropriate volumes/files need to be backed up. If using a Local filesystem, the directory being used for file storage should be backed up. ### Data Lake The path configured for `OUTPUT_DATA_LOCATION` should be backed up. It contains various data files created by the platform. The recommended approach to copy files for your data lake should be used. ### Settings The instance folder of the configurations folder for GGX should be backed up. This does not contain any user data, and can be reconfigured later if needed. This should be done for all the components that are running for GGX - API, App worker, Workers, Jupyter, etc. ## Restoring the systems ### Metadata Database The database dump from the RDBMS system can be reloaded into another database and used. If the `corridor-api db export` method was used to create an SQLite file, the `corridor-api db import` command can be used to import back the SQLite file into the system where the restore is being run. ### File Management The files copied should be kept in the same structure and the user running the GGX process should have permissions to the files on the file management system. ### Data Lake Ideally, the data files should be copied to the same path as the original server where the backup was taken from. ### Settings The setting files can be restored directly and permissions can be set as required. This should be done for all the components that are running for GGX - API, App worker, Workers, Jupyter, etc. --- # Multiple GGX Workers Source: https://docs.genguardx.ai/technology/self-hosting/scaling/concurrency/ Markdown: https://docs.genguardx.ai/technology/self-hosting/scaling/concurrency/index.md Description: Run multiple named GGX workers with separate queues, process files, logs, state databases, and worker-specific configuration in self-hosted environments. GGX provides an option to run multiple workers on the same server, without the workers interfering with each other. The user needs to provide a name for each worker and worker-specific configuration in api_config.py, where each configuration is tied to the worker's name. ## Custom `corridor-worker run` command with a worker name The worker name can be provided with the option `--worker-name` or `-n` - `INSTALL_DIR/venv-api/bin/corridor-worker run --worker-name CUSTOM_` :::note The worker's name needs to have all **capital letters**. Underscores (`_`) can be part of the worker's name. ::: ## Custom worker configurations Any worker-specific configuration is required to be added to the file: - `INSTALL_DIR/instances/INSTANCE_NAME/config/api_config.py` To avoid synchronization issues with other workers running on the same server. The worker configurations have to be prefixed with the worker name provided in the `corridor-worker run` command above. ### Configurations Taking the above `corridor-worker run` command as an example, where the worker name is `CUSTOM_`, the configurations would be: - `CUSTOM_WORKER_QUEUES` - `CUSTOM_WORKER_PIDFILE` - `CUSTOM_WORKER_LOGFILE` - `CUSTOM_WORKER_PROCESSES` - `CUSTOM_CELERY_WORKER_STATE_DB` - `CUSTOM_CELERY_WORKER_HIJACK_ROOT_LOGGER` - `CUSTOM_CELERY_WORKER_REDIRECT_STDOUTS` --- # Scalability and Sizing Source: https://docs.genguardx.ai/technology/self-hosting/scaling/scalability/ Markdown: https://docs.genguardx.ai/technology/self-hosting/scaling/scalability/index.md Description: Plan self-hosted GGX scalability for production throughput, simulation jobs, user analytics, application services, databases, APIs, and worker capacity. If the platform seems to be slow, there may be an infrastructure constraint in terms of resources that could be causing this. This section describes how different aspects of scalability have been kept in mind when designing the platform. The way to think about scalability in the platform is: - Production - Simulation and other jobs - User-initiated analytics - Application ## Production The production scalability again can be thought of in two parts: - Time to execute the artifact-bundle/policy - Throughput for the production orchestration The time taken for the artifact-bundle depends on the complexity of the logic written in the platform. This can be sped up by writing more efficient codes or simplifying the logic. The throughput for the orchestrations depends on the method of deployment of the Production layer, whether it is an HTTP API, serverless architecture, python runtime execution, etc. ## Simulation and other jobs The simulations and other jobs like Comparison, Validation, etc. Use the Celery workers to perform tasks. There are 2 parts which can be considered here: - Number of parallel jobs to run - Cluster nodes available for use per job If the number of parallel jobs is a concern, due to a large job blocking later jobs, etc. - this can be handled by adding more "Spark - Celery Workers" as needed. Note that adding too many Spark - Celery workers could cause resource contention between the workers themselves and may require an increase in cluster nodes. If the spark jobs themselves are slow - then the issue is more likely the cluster nodes available per job. i.e. the number of Spark workers or YARN containers that are allocated per job. To make this faster, more spark nodes need to be provided to make Spark jobs run faster. ## User-initiated analytics The platform's Notebook section allows the user to initiate their own jobs in a free-form manner using Python/PySpark. As these jobs are user-dependent, they can utilize a large number of cluster resources if not used correctly. If this causes issues with the Simulation jobs, it is recommended to use separate YARN queues to manage the resource allocation appropriately. ## Application The Application's scalability is an important factor to consider when the number of concurrent users starts increasing over time. This can impact the usability of the application itself and can be scaled by identifying which portion of the application is causing the bottleneck. The Metadata Database or the API Server could be the cause of issues here. It can be easily rectified by adding more resources to these components or parallelizing them with Database Read Replicas or API Servers appropriately. :::note The API Servers are stateless and can be easily scaled up/down as needed. :::